flâneur

Pretraining progress is mostly coming from data

dwarkesh.com · 2,795 words · saved by 1 readers

Breaking down 6 years of pretraining progress into data vs model improvements

How much of the rapid progress in AI that we’ve seen over the last few years1 has come from data versus model improvements? The answer has big implications for the economics of frontier labs and the pace of future progress. We investigate this question at a relatively small scale, and for pretraining specifically, from 2019 to 2025. During each of those years, a new open model recipe was published which codified that year’s publicly known algorithmic tweaks (for example, improvements in architecture, optimizer, initializations, learning rate schedule, hyperparams, etc). And during each of…

saved by

related reading