How to Think About GPUs | How To Scale Your Model
We love TPUs at Google, but GPUs are great too. This chapter takes a deep dive into the world of NVIDIA GPUs – how each chip works, how they’re networked together, and what that means for LLMs, especially compared to TPUs. This section builds on Chapter 2 and Chapter 5, so you are encouraged to read them first. Jacob Austin† †Google DeepMind Swapnil Patil† Adam Paszke† Reiner Pope* *MatX Aug. 18, 2025 A modern ML GPU (e.g. H100, B200) is basically a bunch of compute cores that specialize in matrix multiplication (called Streaming Multiprocessors or SMs) connected to a stick of fast memory (called HBM). Here’s a diagram: Each SM, like a TPU’s Tensor Core, has a dedicated matrix multiplication core (unfortunately also called a Tensor Core), a vector arithmetic unit (called a Warp Scheduler), and a fast on-chip cache (called SMEM). Unlike a TPU, which has at most 2 independent “Tensor Cores”, a modern GPU has more than 100 SMs (132 on an H100). Each of these SMs is much less powerful than
How to Think About GPUs | How To Scale Your Model How to Think About GPUs Part 12 of How To Scale Your Model ( Part 11: Conclusion | The End ) We love TPUs at Google, but GPUs are great too. This chapter takes a deep dive into the world of GPUs – how each chip works, how they're networked together, and what that means for LLMs, especially compared to TPUs. While there are a multitude of GPU architectures from NVIDIA, AMD, Intel, and others, here we will focus on NVIDIA GPUs. This section builds on Chapter 2 and Chapter 5 , so you are encouraged to read them first. Authors Affiliation Jacob Aus
Explore this link on the map →saved by
related reading
- How to Think About TPUs | How To Scale Your Modeljax-ml.github.io
- How To Scale Your Modeljax-ml.github.io
- All About Rooflines | How To Scale Your Modeljax-ml.github.io
- The Best GPUs for Deep Learning in 2023 — An In-depth Analysistimdettmers.com
- Transformer Inference Arithmetic | kipply's blogkipp.ly
- Inside NVIDIA GPUs: Anatomy of high performance matmul kernels - Aleksa Gordićaleksagordic.com
- TPU Deep Divehenryhmko.github.io
- The Best GPUs for Deep Learning in 2023 — An In-depth Analysistimdettmers.com
- Reiner Pope – The math behind how LLMs are trained and serveddwarkesh.com
- GPU Performance Background User's Guide - NVIDIA Docsdocs.nvidia.com
- The Short Case for Nvidia Stock | YouTube Transcript Optimizeryoutubetranscriptoptimizer.com
- Touching the Elephant - TPUs | Consider the Bulldogconsiderthebulldog.com