Pretraining progress is mostly coming from data
dwarkesh.com · 2,795 words · saved by 1 readers
Breaking down 6 years of pretraining progress into data vs model improvements
How much of the rapid progress in AI that we’ve seen over the last few years1 has come from data versus model improvements? The answer has big implications for the economics of frontier labs and the pace of future progress. We investigate this question at a relatively small scale, and for pretraining specifically, from 2019 to 2025. During each of those years, a new open model recipe was published which codified that year’s publicly known algorithmic tweaks (for example, improvements in architecture, optimizer, initializations, learning rate schedule, hyperparams, etc). And during each of…
saved by
related reading
- Scaling is subtler than it seemsberen.io
- frontier model training methodologies | Alex Wa's Blogdjdumpling.github.io
- [2509.14786] Pre-training under infinite computearxiv.org
- Most Algorithmic Progress is Data Progressberen.io
- The Extreme Inefficiency of RL for Frontier Models - Toby Ordtobyord.com
- >10x More Efficient Pretraining — Magicmagic.dev
- AI progress is about to speed up | Epoch AIepoch.ai
- There Are No New Ideas in AI… Only New Datasetsblog.jxmo.io
- I. From GPT-4 to AGI: Counting the OOMs - SITUATIONAL AWARENESSsituational-awareness.ai
- MAI-Thinking-1: Building a Hill-Climbing Machinemicrosoft.ai
- [2509.14786] Pre-training under infinite computearxiv.org
- The least understood driver of AI progress | Epoch AIepoch.ai