MatX: High-throughput chips for LLMs
matx.com · 172 words · saved by 5 readers
We make the best chips physically possible for the large model needs of frontier labs.
Our goal is to make the best chips physically possible for the large model needs of frontier labs. The MatX One chip delivers higher throughput than any announced product while also matching the best latencies of any products. For training and prefill, it excels on FLOPS; for decode and RL it excels on latency, FLOPS, and long-context support. What we offer The highest FLOPS/mm2. Weights are typically in SRAM, for low latency. Allows >2000 output tokens/second for large 100-layer MoE models. KVs are typically in HBM, to support long context well. The most scale-up interconnect of any…
saved by
related reading
- MatX One and our Series B | MatXmatx.com
- How To Scale Your Modeljax-ml.github.io
- An Interview with MatX CEO Reiner Pope About LLM Chipschipstrat.com
- Reiner Pope of MatX on accelerating AI with transformer-optimized chipscheekypint.substack.com
- LLM Engineer's Almanac - Advisormodal.com
- Reiner Pope – The math behind how LLMs are trained and serveddwarkesh.com
- Scalable MatMul-free Language Modelingarxiv.org
- LLM Inference Performance Engineering: Best Practices | Databricks Blogdatabricks.com
- How to Deploy Your Modelhtdym.sailresearch.com
- How is LLaMa.cpp possible?finbarr.ca
- Look Ma, No Bubbles! Designing a Low-Latency Megakernel for Llama-1B · Hazy Researchhazyresearch.stanford.edu
- Big Boss (@0xBADB01E) on Xx.com