Modular: The world's fastest unified matrix multiplication
modular.com · 3,269 words · saved by 1 readers
We are building a next-generation AI developer platform for the world. Read our latest post on how The world's fastest unified matrix multiplication
Modular: The world's fastest unified matrix multiplication Qualcomm to Acquire Modular. Read More → April 20, 2023 The world's fastest unified matrix multiplication Abdul Dakkak Chad Jarvis Eric Johnson Hengjie Wang Ian Tramble Engineering "Matmul", a microcosm of AI performance In our previous blog post , we described why AI needs to solve its compute fragmentation problem to reach its full potential and how matrix multiplication ("matmul") exemplifies why this remains an unsolved problem. In this post, we describe Modular’s approach to solving this problem and its game-changing benefits, inc
related reading
- Modular: AI’s compute fragmentation: what matrix multiplication teaches usmodular.com
- Inside NVIDIA GPUs: Anatomy of high performance matmul kernels - Aleksa Gordićaleksagordic.com
- Modular: Inference from Kernel to Cloudmodular.com
- How to Optimize a CUDA Matmul Kernel for cuBLAS-like Performance: a Worklogsiboehm.com
- Transformer Inference Arithmetic | kipply's blogkipp.ly
- Strangely, Matrix Multiplications on GPUs Run Faster When Given "Predictable" Data! [short]thonking.ai
- MatX: High-throughput chips for LLMsmatx.com
- 1.5x faster MoE training with custom MXFP8 kernels · Cursorcursor.com
- Tiny TPUtinytpu.com
- Scalable MatMul-free Language Modelingarxiv.org
- My picture of the present in AI — LessWronglesswrong.com
- Modern GPU Programming For MLSys — Modern GPU Programming For MLSysmlc.ai