Trust me bro, just one more RL scale up, this one will be the real scale up with the good environments, the actually legit one, trust me bro
I've recently written about how I've updated against seeing substantially faster than trend AI progress due to quickly massively scaling up RL on agentic software engineering. One response I've heard is something like: RL scale-ups so far have used very crappy environments due to difficulty quickly sourcing enough decent (or even high quality) environments. Thus, once AI companies manage to get their hands on actually good RL environments (which could happen pretty quickly), performance will increase a bunch. Another way to put this response is that AI companies haven't actually done a good job scaling up RL—they've scaled up the compute, but with low quality data—and once they actually do the RL scale up for real this time, there will be a big jump in AI capabilities (which yields substantially above trend progress). I'm skeptical of this argument because I think that ongoing improvements to RL environments are already priced into the existing trend: I expect that a substantial part o
I've recently written about how I've updated against seeing substantially faster than trend AI progress due to quickly massively scaling up RL on agentic software engineering. One response I've heard is something like: RL scale-ups so far have used very crappy environments due to difficulty quickly sourcing enough decent (or even high quality) environments. Thus, once AI companies manage to get their hands on actually good RL environments (which could happen pretty quickly), performance will increase a bunch. Another way to put this response is that AI companies haven't actually done a…
related reading
- Trust me bro, just one more RL scale up, this one will be the real scale up with the good environments, the actually legit one, trust me bro — AI Alignment Forumalignmentforum.org
- The upcoming GPT-3 moment for RL | Mechanize, Inc.mechanize.work
- Will scaling work?dwarkeshpatel.com
- RL Environments and RL for Science: Data Foundries and Multi-Agent Architecturesnewsletter.semianalysis.com
- AI progress is about to speed up | Epoch AIepoch.ai
- Akash Bajwa on X: "RL Environments with Scale AI" / Xx.com
- Scaling Reinforcement Learning: Environments, Reward Hacking, Agents, Scaling Datasemianalysis.com
- What's going on with AI progress and trends? (As of 5/2025)redwoodresearch.substack.com
- AI in 2025: gestalt — LessWronglesswrong.com
- The least understood driver of AI progress | Epoch AIepoch.ai
- The Extreme Inefficiency of RL for Frontier Models - Toby Ordtobyord.com
- Dario Amodei — "We are near the end of the exponential"dwarkesh.com