How To Scale Your Model
Training LLMs often feels like alchemy, but understanding and optimizing the performance of your models doesn't have to. This book aims to demystify the science of scaling language models: how TPUs (and GPUs) work and how they communicate with each other, how LLMs run on real hardware, and how to parallelize your models during training and inference so they run efficiently at massive scale. If you've ever wondered
How To Scale Your Model How to Scale Your Model A Systems View of LLMs on TPUs (Part 0: Intro | Part 1: Rooflines ) Training LLMs often feels like alchemy, but understanding and optimizing the performance of your models doesn't have to. This book aims to demystify the science of scaling language models: how TPUs (and GPUs) work and how they communicate with each other, how LLMs run on real hardware, and how to parallelize your models during training and inference so they run efficiently at massive scale. If you've ever wondered "how expensive should this LLM be to train" or "how much memory do
Explore this link on the map →saved by
- Tazik Sh
- Karan MJ
- Tasha Pais
- Ratan Kaliani
- Nikki Guo
- Benedict Neo
- Asher P
- Sarah Pan
- Jianmin Chen
- Aaron Pham
- Jirat C
- Sahil Jain
related reading
- Transformer Inference Arithmetic | kipply's blogkipp.ly
- All About Rooflines | How To Scale Your Modeljax-ml.github.io
- Reiner Pope – The math behind how LLMs are trained and serveddwarkesh.com
- Transformer Math 101 | EleutherAI Blogblog.eleuther.ai
- How to Think About GPUs | How To Scale Your Modeljax-ml.github.io
- TPU Deep Divehenryhmko.github.io
- How to Think About TPUs | How To Scale Your Modeljax-ml.github.io
- Transformers Inference Optimization Toolset | AstraBlogastralord.github.io
- Demystify Transformers: A Guide to Scaling Laws | by Yu-Cheng Tsai | Sage Ai | Mediummedium.com
- Bits, FLOPS, and Watts: A Systems-Level Perspective of Scaling LLMs — Part 1 | by Asheesh Goja | Mediummedium.com
- Pipeline-Parallelism: Distributed Training via Model Partitioningsiboehm.com
- LLM Inference Performance Engineering: Best Practices | Databricks Blogdatabricks.com