Data movement bottlenecks to large-scale model training: Scaling past 1e28 FLOP | Epoch AI
Data movement bottlenecks limit LLM scaling beyond 2e28 FLOP, with a “latency wall” at 2e31 FLOP. We may hit these in ~3 years. Aggressive batch size scaling could potentially overcome these limits.
Data movement bottlenecks to large-scale model training: Scaling past 1e28 FLOP | Epoch AI Introduction Over the past five years, the performance of large language models (LLMs) has improved dramatically, driven largely by rapid scaling in training compute budgets to handle larger models and training datasets. Our own estimates suggest that the training compute used by frontier AI models has grown by 4-5 times every year from 2010 to 2024. This rapid pace of scaling far outpaces Moore’s law, and sustaining it has required scaling along three dimensions: First, making training runs last longer;
Explore this link on the map →related reading
- How To Scale Your Modeljax-ml.github.io
- Reiner Pope – The math behind how LLMs are trained and serveddwarkesh.com
- Transformer Inference Arithmetic | kipply's blogkipp.ly
- Can AI scaling continue through 2030? | Epoch AIepoch.ai
- The Scaling Hypothesis · Gwern.netgwern.net
- Bits, FLOPS, and Watts: A Systems-Level Perspective of Scaling LLMs — Part 1 | by Asheesh Goja | Mediummedium.com
- Fermi estimate of future training runsdanieldewey.net
- Demystify Transformers: A Guide to Scaling Laws | by Yu-Cheng Tsai | Sage Ai | Mediummedium.com
- All About Rooflines | How To Scale Your Modeljax-ml.github.io
- LLM Inference Performance Engineering: Best Practices | Databricks Blogdatabricks.com
- Multi-Datacenter Training: OpenAI's Ambitious Plan To Beat Google's Infrastructuresemianalysis.com
- Transformer Math 101 | EleutherAI Blogblog.eleuther.ai