flâneur — a map of the web's best reading

Data movement bottlenecks to large-scale model training: Scaling past 1e28 FLOP | Epoch AI

epoch.ai · 3,665 words · saved by 1 readers

Data movement bottlenecks limit LLM scaling beyond 2e28 FLOP, with a “latency wall” at 2e31 FLOP. We may hit these in ~3 years. Aggressive batch size scaling could potentially overcome these limits.

Data movement bottlenecks to large-scale model training: Scaling past 1e28 FLOP | Epoch AI Introduction Over the past five years, the performance of large language models (LLMs) has improved dramatically, driven largely by rapid scaling in training compute budgets to handle larger models and training datasets. Our own estimates suggest that the training compute used by frontier AI models has grown by 4-5 times every year from 2010 to 2024. This rapid pace of scaling far outpaces Moore’s law, and sustaining it has required scaling along three dimensions: First, making training runs last longer;

Explore this link on the map →

related reading