✳flâneur — a map of the web's best reading
RL Scaling Laws for LLMs - by Cameron R. Wolfe, Ph.D.
cameronrwolfe.substack.com · 12,775 words · saved by 1 readers
How scaling laws have evolved from pretraining to reinforcement learning...
RL Scaling Laws for LLMs How scaling laws have evolved from pretraining to reinforcement learning... Cameron R. Wolfe, Ph.D. Apr 20, 2026 129 17 Share (from [1, 2, 3]) Scaling is one of the most impactful concepts in the history of AI research. For large language models (LLMs), scaling has mostly been studied in the context of pretraining, where rigorous scaling laws have allowed us to clearly define the relationship between compute and performance. Inspired by these predictable trends, the LLM research community has empirically validated pretraining scaling laws across several orders of magni
Explore this link on the map →saved by
related reading
- State of RL for reasoning LLMs | A. Weersaweers.de
- Xiuyu Li on X: "RL Interview Questions 2026" / Xx.com
- A vision researcher’s guide to some RL stuff: PPO & GRPO - Yuge (Jimmy) Shiyugeten.github.io
- IsoCompute Playbook: Optimally Scaling Sampling Compute for RL Training of LLMscompute-optimal-rl-llm-scaling.github.io
- LoRA Without Regret - Thinking Machines Labthinkingmachines.ai
- [2510.13786] The Art of Scaling Reinforcement Learning Compute for LLMsarxiv.org
- The State of Reinforcement Learning for LLM Reasoningmagazine.sebastianraschka.com
- Composer2.pdfcursor.com
- LLM Resourcesforrestbicker.com
- From REINFORCE to Dr. GRPOlancelqf.github.io
- PPO for LLMs: A Guide for Normal Peoplecameronrwolfe.substack.com
- Understanding R1-Zero-Like Training: A Critical Perspectivearxiv.org