The Extreme Inefficiency of RL for Frontier Models — Toby Ord
The new scaling paradigm for AI reduces the amount of information a model can learn from per hour of training by a factor of 1,000 to 1,000,000. I explore what this means and its implications for scaling. The last year has seen a massive shift in how leading AI models are trained. 2018–2023 was the era of pre-training scaling. LLMs were primarily trained by next-token prediction (also known as pre-training). Much of OpenAI’s progress from GPT-1 to GPT-4, came from scaling up the amount of pre-training by a factor of 1,000,000. New capabilities were unlocked not through scientific breakthroughs, but through doing more-or-less the same thing at ever-larger scales. Everyone was talking about the success of scaling, from AI labs to venture capitalists to policy makers. However, there’s been markedly little progress in scaling up this kind of training since (GPT-4.5 added one more factor of 10, but was then quietly retired). Instead, there has been a shift to taking one of these pre-trained
The Extreme Inefficiency of RL for Frontier Models September 19, 2025 Toby Ord The new scaling paradigm for AI reduces the amount of information a model can learn from per hour of training by a factor of 1,000 to 1,000,000. I explore what this means and its implications for scaling. The last year has seen a massive shift in how leading AI models are trained. 2018–2023 was the era of pre-training scaling. LLMs were primarily trained by next-token prediction (also known as pre-training). Much of OpenAI’s progress from GPT-1 to GPT-4, came from scaling up the amount of pre-training by a factor of
Explore this link on the map →related reading
- AI in 2025: gestalt — LessWronglesswrong.com
- RL is even more information inefficient than you thoughtdwarkesh.com
- The Scaling Hypothesis · Gwern.netgwern.net
- How to scale RL to 10^26 FLOPs - by Jack Morrisblog.jxmo.io
- RL Scaling Laws for LLMs - by Cameron R. Wolfe, Ph.D.cameronrwolfe.substack.com
- How Well Does RL Scale? - Toby Ordtobyord.com
- Andrej Karpathy — AGI is still a decade awaydwarkesh.com
- How can LLM RL Work Despite Information-Theoretic Inefficiencyberen.io
- State of RL for reasoning LLMs | A. Weersaweers.de
- AI progress is about to speed up | Epoch AIepoch.ai
- Dario Amodei — "We are near the end of the exponential"dwarkesh.com
- What I've Learned About AI in the Past Two Months.sheracaolity.ghost.io