✳flâneur — a map of the web's best reading
Modular: The world's fastest unified matrix multiplication
modular.com · 3,269 words · saved by 1 readers
We are building a next-generation AI developer platform for the world. Read our latest post on how The world's fastest unified matrix multiplication
Modular: The world's fastest unified matrix multiplication Qualcomm to Acquire Modular. Read More → April 20, 2023 The world's fastest unified matrix multiplication Abdul Dakkak Chad Jarvis Eric Johnson Hengjie Wang Ian Tramble Engineering "Matmul", a microcosm of AI performance In our previous blog post , we described why AI needs to solve its compute fragmentation problem to reach its full potential and how matrix multiplication ("matmul") exemplifies why this remains an unsolved problem. In this post, we describe Modular’s approach to solving this problem and its game-changing benefits, inc
Explore this link on the map →related reading
- Modular: AI’s compute fragmentation: what matrix multiplication teaches usmodular.com
- Inside NVIDIA GPUs: Anatomy of high performance matmul kernels - Aleksa Gordićaleksagordic.com
- Modular: Inference from Kernel to Cloudmodular.com
- How to Optimize a CUDA Matmul Kernel for cuBLAS-like Performance: a Worklogsiboehm.com
- Transformer Inference Arithmetic | kipply's blogkipp.ly
- 1.5x faster MoE training with custom MXFP8 kernels · Cursorcursor.com
- Strangely, Matrix Multiplications on GPUs Run Faster When Given "Predictable" Data! [short]thonking.ai
- Tiny TPUtinytpu.com
- My picture of the present in AI — LessWronglesswrong.com
- [2410.20399] ThunderKittens: Simple, Fast, and Adorable AI Kernelsarxiv.org
- The Short Case for Nvidia Stock | YouTube Transcript Optimizeryoutubetranscriptoptimizer.com
- Understanding Matrix Multiplication on a Weight-Stationary Systolic Architecture | Telesenstelesens.co