How to Think About TPUs | How To Scale Your Model
This section is all about how TPUs work, how they're networked together to enable multi-chip training and inference, and how this affects the performance of our favorite algorithms. There's even some good stuff for GPU users too!
How to Think About TPUs | How To Scale Your Model How to Think About TPUs Part 2 of How To Scale Your Model ( Part 1: Rooflines | Part 3: Sharding ) This section is all about how TPUs work, how they're networked together to enable multi-chip training and inference, and how this affects the performance of our favorite algorithms. There's even some good stuff for GPU users too! Authors Affiliation Jacob Austin Google DeepMind Sholto Douglas Roy Frostig Anselm Levskaya Charlie Chen Sharad Vikram Federico Lebron Peter Choy Vinay Ramasesh Albert Webson Reiner Pope * Published Feb. 4, 2025 You might
Explore this link on the map →saved by
related reading
- All About Rooflines | How To Scale Your Modeljax-ml.github.io
- How to Think About GPUs | How To Scale Your Modeljax-ml.github.io
- How To Scale Your Modeljax-ml.github.io
- TPU Deep Divehenryhmko.github.io
- Touching the Elephant - TPUs | Consider the Bulldogconsiderthebulldog.com
- Inside NVIDIA GPUs: Anatomy of high performance matmul kernels - Aleksa Gordićaleksagordic.com
- Tiny TPUtinytpu.com
- CVPR2023_eff_tutorial_molchanov.pdfnvlabs.github.io
- Transformer Inference Arithmetic | kipply's blogkipp.ly
- Making Deep Learning go Brrrr From First Principleshorace.io
- NVIDIA Tensor Core Evolution: From Volta To Blackwellnewsletter.semianalysis.com
- 1.5x faster MoE training with custom MXFP8 kernels · Cursorcursor.com