Understanding Matrix Multiplication on a Weight-Stationary Systolic Architecture | Telesens
If you follow the hardware for deep learning space, you may have heard of the term “systolic array”. A 2D systolic array forms the heart of the Matrix Multiplier Unit (MXU) on the Google TPU and the new deep learning FPGAs from Xilinx. If you are a computer architecture expert, then you know what systolic arrays are and perhaps even implemented a convolution or matrix multiplication on a systolic array in grad school. If you are like me – generally well versed in tech, but not a computer architecture expert, then you are probably a bit lost. In this post, I’ll explain what systolic arrays are and how a 1D correlation operation can be implemented using a systolic array. I’ll then show an animation of how the multiplication of two matrices is implemented on a systolic array, which will help you understand the trade-offs made in systolic architectures. The operation of the MXU on a TPU is identical to the data flow shown in the animation. To provide a concrete example of the ideas discus
If you follow the hardware for deep learning space, you may have heard of the term “systolic array”. A 2D systolic array forms the heart of the Matrix Multiplier Unit (MXU) on the Google TPU and the new deep learning FPGAs from Xilinx. If you are a computer architecture expert, then you know what systolic arrays are and perhaps even implemented a convolution or matrix multiplication on a systolic array in grad school. If you are like me – generally well versed in tech, but not a computer architecture expert, then you are probably a bit lost. In this post, I’ll explain w
Explore this link on the map →saved by
related reading
- Tiny TPUtinytpu.com
- Inside NVIDIA GPUs: Anatomy of high performance matmul kernels - Aleksa Gordićaleksagordic.com
- TPU Deep Divehenryhmko.github.io
- How To Scale Your Modeljax-ml.github.io
- How to Think About TPUs | How To Scale Your Modeljax-ml.github.io
- All About Rooflines | How To Scale Your Modeljax-ml.github.io
- Touching the Elephant - TPUs | Consider the Bulldogconsiderthebulldog.com
- Making Deep Learning go Brrrr From First Principleshorace.io
- How to Optimize a CUDA Matmul Kernel for cuBLAS-like Performance: a Worklogsiboehm.com
- Transformer Inference Arithmetic | kipply's blogkipp.ly
- The Best GPUs for Deep Learning in 2023 — An In-depth Analysistimdettmers.com
- GPU Performance Background User's Guide - NVIDIA Docsdocs.nvidia.com