How Well Does RL Scale? — Toby Ord
The current era of improving AI capabilities using Reinforcement Learning (from verifiable rewards) involves two key types of scaling: Scaling the amount of compute used for RL during training Scaling the amount of compute used for inference during deployment We can see (1) as training the AI in more effective reasoning techniques and (2) as allowing the model to think for longer. I’ll call the first RL-scaling, and the second inference-scaling. Both new kinds of scaling were present all the way back in OpenAI’s announcement of their first reasoning model, o1, when they showed this famous chart: I’ve previously shown that in the initial move from a base-model to a reasoning model, most of the performance gain came from unlocking the inference-scaling. The RL training did provide a notable boost to performance, even holding the number of tokens in the chain of thought fixed. You can see this RL boost in the chart below as the small blue arrow on the left that takes the base model up to
How Well Does RL Scale? October 20, 2025 Toby Ord The current era of improving AI capabilities using Reinforcement Learning (from verifiable rewards) involves two key types of scaling: Scaling the amount of compute used for RL during training Scaling the amount of compute used for inference during deployment We can see (1) as training the AI in more effective reasoning techniques and (2) as allowing the model to think for longer. I’ll call the first RL-scaling , and the second inference-scaling . Both new kinds of scaling were present all the way back in OpenAI’s announcement of their first re
Explore this link on the map →related reading
- The Scaling Hypothesis · Gwern.netgwern.net
- RL Scaling Laws for LLMs - by Cameron R. Wolfe, Ph.D.cameronrwolfe.substack.com
- [2510.13786] The Art of Scaling Reinforcement Learning Compute for LLMsarxiv.org
- How to scale RL to 10^26 FLOPs - by Jack Morrisblog.jxmo.io
- IsoCompute Playbook: Optimally Scaling Sampling Compute for RL Training of LLMscompute-optimal-rl-llm-scaling.github.io
- The Extreme Inefficiency of RL for Frontier Models - Toby Ordtobyord.com
- AI in 2025: gestalt — LessWronglesswrong.com
- Inference Scaling Reshapes AI Governance - Toby Ordtobyord.com
- AI progress is about to speed up | Epoch AIepoch.ai
- Distinguish between inference scaling and "larger tasks use more compute" — AI Alignment Forumalignmentforum.org
- State of RL for reasoning LLMs | A. Weersaweers.de
- Scaling Reinforcement Learning: Environments, Reward Hacking, Agents, Scaling Datasemianalysis.com