RL Scaling Laws for LLMs - by Cameron R. Wolfe, Ph.D.
cameronrwolfe.substack.com · 12,775 words · saved by 1 readers
How scaling laws have evolved from pretraining to reinforcement learning...
RL Scaling Laws for LLMs How scaling laws have evolved from pretraining to reinforcement learning... Cameron R. Wolfe, Ph.D. Apr 20, 2026 129 17 Share (from [1, 2, 3]) Scaling is one of the most impactful concepts in the history of AI research. For large language models (LLMs), scaling has mostly been studied in the context of pretraining, where rigorous scaling laws have allowed us to clearly define the relationship between compute and performance. Inspired by these predictable trends, the LLM research community has empirically validated pretraining scaling laws across several orders of magni
saved by
related reading
- State of RL for reasoning LLMs | A. Weersaweers.de
- Xiuyu Li on X: "RL Interview Questions 2026" / Xx.com
- How can LLM RL Work Despite Information-Theoretic Inefficiencyberen.io
- Scaling is subtler than it seemsberen.io
- A vision researcher’s guide to some RL stuff: PPO & GRPO - Yuge (Jimmy) Shiyugeten.github.io
- [2510.13786] The Art of Scaling Reinforcement Learning Compute for LLMsarxiv.org
- IsoCompute Playbook: Optimally Scaling Sampling Compute for RL Training of LLMscompute-optimal-rl-llm-scaling.github.io
- GRPO++: Tricks for Making RL Actually Workcameronrwolfe.substack.com
- Scaling Laws, Carefully | Lil'Loglilianweng.github.io
- The Extreme Inefficiency of RL for Frontier Models - Toby Ordtobyord.com
- LoRA Without Regret - Thinking Machines Labthinkingmachines.ai
- The upcoming GPT-3 moment for RL | Mechanize, Inc.mechanize.work