An audio version of my blog post, Thoughts on AI progress (Dec 2025)
I’m confused why some people have short timelines and at the same time are bullish on the current scale up of reinforcement learning atop LLMs. If we’re actually close to a human-like learner, this whole approach of training on verifiable outcomes is doomed. Currently the labs are trying to bake in a bunch of skills into these models through “mid-training” - there’s an entire supply chain of companies building RL environments which teach the model how to navigate a web browser or use Excel to write financial models. Either these models will soon learn on the job in a self directed way - making all this pre-baking pointless - or they won’t - which means AGI is not imminent. Humans don’t have to go through a special training phase where they need to rehearse every single piece of software they might ever need to use. Beren Millidge made interesting points about this in a recent blog post: When we see frontier models improving at various benchmarks we should think not just of increased sc
I’m confused why some people have short timelines and at the same time are bullish on the current scale up of reinforcement learning atop LLMs. If we’re actually close to a human-like learner, this whole approach of training on verifiable outcomes is doomed. Currently the labs are trying to bake in a bunch of skills into these models through “mid-training” - there’s an entire supply chain of companies building RL environments which teach the model how to navigate a web browser or use Excel to write financial models. Either these models will soon learn on the job in a self directed way -…
saved by
related reading
- Thoughts on AI progress (Dec 2025)dwarkesh.com
- Thoughts on AI progress (Dec 2025)substack.com
- Thoughts on AI progress (Dec 2025) - by Dwarkesh Pateldwarkesh.com
- Andrej Karpathy — AGI is still a decade awaydwarkesh.com
- AI in 2025: gestalt — LessWronglesswrong.com
- Will scaling work?dwarkeshpatel.com
- Questions about the Future of AI - by Dwarkesh Pateldwarkesh.com
- The Extreme Inefficiency of RL for Frontier Models - Toby Ordtobyord.com
- 2025 LLM Year in Review – karpathykarpathy.bearblog.dev
- Dario Amodei — "We are near the end of the exponential"dwarkesh.com
- Why I don’t think AGI is right around the cornerdwarkesh.com
- The upcoming GPT-3 moment for RL | Mechanize, Inc.mechanize.work