How Does Claude 4 Think? — Sholto Douglas & Trenton Bricken
New episode with my good friends Sholto Douglas & Trenton Bricken. Sholto focuses on scaling RL and Trenton researches mechanistic interpretability, both at Anthropic. We talk through what’s changed in the last year of AI research; the new RL regime and how far it can scale; how to trace a model’s thoughts; and how countries, workers, and students should prepare for AGI. See you next year for v3. Here’s last year’s episode, btw. Enjoy! Watch on YouTube; listen on Apple Podcasts or Spotify. WorkOS ensures that AI companies like OpenAI and Anthropic don't have to spend engineering time building enterprise features like access controls or SSO. It’s not that they don't need these features; it's just that WorkOS gives them battle-tested APIs that they can use for auth, provisioning, and more. Start building today at workos.com. Scale is building the infrastructure for safer, smarter AI. Scale’s Data Foundry gives major AI labs access to high-quality data to fuel post-training, while their p
Playback speed × Share post Share post at current time Share from 0:00 0:00 / Generate transcript A transcript unlocks clips, previews, and editing. 116 2 7 Is RL + LLMs enough for AGI? — Sholto Douglas & Trenton Bricken Scaling reinforcement learning, tracing circuits, and the path to fully autonomous agents Dwarkesh Patel May 22, 2025 116 2 7 Share Transcript New episode with my good friends Sholto Douglas & Trenton Bricken . Sholto focuses on scaling RL and Trenton researches mechanistic interpretability, both at Anthropic. We talk through what’s changed in the last year of AI research; the
Explore this link on the map →related reading
- Tracing the thoughts of a large language model \ Anthropicanthropic.com
- AI in 2025: gestalt — LessWronglesswrong.com
- DeepSeek-R1arxiv.org
- Andrej Karpathy — AGI is still a decade awaydwarkesh.com
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- Mapping the Mind of a Large Language Model \ Anthropicanthropic.com
- Tracing the Thoughts of a Large Language Model — LessWronglesswrong.com
- Dario Amodei — "We are near the end of the exponential"dwarkesh.com
- Automated Alignment Researchers: Using large language models to scale scalable oversight \ Anthropicanthropic.com
- Questions about the Future of AI - by Dwarkesh Pateldwarkesh.com
- How Can Interpretability Researchers Help AGI Go Well? — AI Alignment Forumalignmentforum.org
- The State of Reinforcement Learning for LLM Reasoningmagazine.sebastianraschka.com