Scaling Reinforcement Learning: Environments, Reward Hacking, Agents, Scaling Data – SemiAnalysis
The test time scaling paradigm is thriving. Reasoning models continue to rapidly improve, and are becoming more effective and affordable. Evaluations measuring real world software engineering tasks, like SWE-Bench, are seeing higher scores at cheaper costs. Below is a chart showing how models are both getting cheaper and better. Reinforcement learning (RL) is the reason for this progress. We covered this in a previous report, outlining how RL has unlocked the ability for models to do reasoning through generating a Chain of Thought (CoT). We expect this paradigm to continue. Other than just CoT innovations, models are coherent (think) for longer, which unlock agentic capabilities. Tool use such as search, utilizing python for calculations and other capabilities are downstream from the ability to plan, reason, and operate for long amounts of time. Better reasoning gives models more time to “think”, and thus emerge from simple chatbots to planners. This, in turn, enables more coherent age
Scaling Reinforcement Learning: Environments, Reward Hacking, Agents, Scaling Data Infrastructure Bottlenecks and Changes, Distillation, Data is a Moat, Recursive Self Improvement, o4 and o5 RL Training, China Accelerator Production Dylan Patel and AJ Jun 08, 2025 ∙ Paid 20 2 Share The test time scaling paradigm is thriving. Reasoning models continue to rapidly improve, and are becoming more effective and affordable. Evaluations measuring real world software engineering tasks, like SWE-Bench, are seeing higher scores at cheaper costs. Below is a chart showing how models are both getting cheape
Explore this link on the map →related reading
- Reward Hacking in Reinforcement Learning | Lil'Loglilianweng.github.io
- Composer2.pdfcursor.com
- Akash Bajwa on X: "RL Environments with Scale AI" / Xx.com
- RL Environments and RL for Science: Data Foundries and Multi-Agent Architecturesnewsletter.semianalysis.com
- Trust me bro, just one more RL scale up, this one will be the real scale up with the good environments, the actually legit one, trust me bro — AI Alignment Forumalignmentforum.org
- How to scale RL to 10^26 FLOPs - by Jack Morrisblog.jxmo.io
- RLHF: Reinforcement Learning from Human Feedbackhuyenchip.com
- State of RL for reasoning LLMs | A. Weersaweers.de
- IsoCompute Playbook: Optimally Scaling Sampling Compute for RL Training of LLMscompute-optimal-rl-llm-scaling.github.io
- How Well Does RL Scale? - Toby Ordtobyord.com
- Dario Amodei — "We are near the end of the exponential"dwarkesh.com
- A Taxonomy of RL Environments for LLM Agentsleehanchung.github.io