MatX One and our Series B | MatX
We’re building an LLM chip that delivers much higher throughput than any other chip while also achieving the lowest latency. We call it the MatX One. The MatX One chip is based on a splittable systolic array, which has the energy and area efficiency that large systolic arrays are famous for, while also getting high utilization on smaller matrices with flexible shapes. The chip combines the low latency of SRAM-first designs with the long-context support of HBM. These elements, plus a fresh take on numerics, deliver higher throughput on LLMs than any announced system, while simultaneously matching the latency of SRAM-first designs. Higher throughput and lower latency give you smarter and faster models for your subscription dollar. We’ve raised a $500M Series B to wrap up development and quickly scale manufacturing, with tapeout in under a year. The round was led by Jane Street, one of the most tech-savvy Wall Street firms, and Situational Awareness LP, whose founder Leopold Aschenbrenner
February 24, 2026 Reiner Pope We’re building an LLM chip that delivers much higher throughput than any other chip while also achieving the lowest latency. We call it the MatX One. The MatX One chip is based on a splittable systolic array, which has the energy and area efficiency that large systolic arrays are famous for, while also getting high utilization on smaller matrices with flexible shapes. The chip combines the low latency of SRAM-first designs with the long-context support of HBM. These elements, plus a fresh take on numerics, deliver higher throughput on LLMs than any announced syste
Explore this link on the map →related reading
- An Interview with MatX CEO Reiner Pope About LLM Chipschipstrat.com
- MatX: High-throughput chips for LLMsmatx.com
- Reiner Pope of MatX on accelerating AI with transformer-optimized chipscheekypint.substack.com
- How To Scale Your Modeljax-ml.github.io
- How to Land a Frontier Lab Jobvladfeinberg.com
- Reiner Pope – The math behind how LLMs are trained and serveddwarkesh.com
- Look Ma, No Bubbles! Designing a Low-Latency Megakernel for Llama-1B · Hazy Researchhazyresearch.stanford.edu
- LLM Inference Performance Engineering: Best Practices | Databricks Blogdatabricks.com
- ⭐️ Fast LLM Inference From Scratchandrewkchan.dev
- How is LLaMa.cpp possible?finbarr.ca
- Inside NVIDIA GPUs: Anatomy of high performance matmul kernels - Aleksa Gordićaleksagordic.com
- LLM Engineer's Almanac - Workloads | Modalmodal.com