flâneur — a map of the web's best reading

Scaling Reinforcement Learning: Environments, Reward Hacking, Agents, Scaling Data – SemiAnalysis

semianalysis.com · 6,800 words · saved by 1 readers

The test time scaling paradigm is thriving. Reasoning models continue to rapidly improve, and are becoming more effective and affordable. Evaluations measuring real world software engineering tasks, like SWE-Bench, are seeing higher scores at cheaper costs. Below is a chart showing how models are both getting cheaper and better. Reinforcement learning (RL) is the reason for this progress. We covered this in a previous report, outlining how RL has unlocked the ability for models to do reasoning through generating a Chain of Thought (CoT). We expect this paradigm to continue. Other than just CoT innovations, models are coherent (think) for longer, which unlock agentic capabilities. Tool use such as search, utilizing python for calculations and other capabilities are downstream from the ability to plan, reason, and operate for long amounts of time. Better reasoning gives models more time to “think”, and thus emerge from simple chatbots to planners. This, in turn, enables more coherent age

Scaling Reinforcement Learning: Environments, Reward Hacking, Agents, Scaling Data Infrastructure Bottlenecks and Changes, Distillation, Data is a Moat, Recursive Self Improvement, o4 and o5 RL Training, China Accelerator Production Dylan Patel and AJ Jun 08, 2025 ∙ Paid 20 2 Share The test time scaling paradigm is thriving. Reasoning models continue to rapidly improve, and are becoming more effective and affordable. Evaluations measuring real world software engineering tasks, like SWE-Bench, are seeing higher scores at cheaper costs. Below is a chart showing how models are both getting cheape

Explore this link on the map →

related reading