flâneur — a map of the web's best reading

How Does Claude 4 Think? — Sholto Douglas & Trenton Bricken

dwarkesh.com · 25,507 words · saved by 1 readers

New episode with my good friends Sholto Douglas & Trenton Bricken. Sholto focuses on scaling RL and Trenton researches mechanistic interpretability, both at Anthropic. We talk through what’s changed in the last year of AI research; the new RL regime and how far it can scale; how to trace a model’s thoughts; and how countries, workers, and students should prepare for AGI. See you next year for v3. Here’s last year’s episode, btw. Enjoy! Watch on YouTube; listen on Apple Podcasts or Spotify. WorkOS ensures that AI companies like OpenAI and Anthropic don't have to spend engineering time building enterprise features like access controls or SSO. It’s not that they don't need these features; it's just that WorkOS gives them battle-tested APIs that they can use for auth, provisioning, and more. Start building today at workos.com. Scale is building the infrastructure for safer, smarter AI. Scale’s Data Foundry gives major AI labs access to high-quality data to fuel post-training, while their p

Playback speed × Share post Share post at current time Share from 0:00 0:00 / Generate transcript A transcript unlocks clips, previews, and editing. 116 2 7 Is RL + LLMs enough for AGI? — Sholto Douglas & Trenton Bricken Scaling reinforcement learning, tracing circuits, and the path to fully autonomous agents Dwarkesh Patel May 22, 2025 116 2 7 Share Transcript New episode with my good friends Sholto Douglas & Trenton Bricken . Sholto focuses on scaling RL and Trenton researches mechanistic interpretability, both at Anthropic. We talk through what’s changed in the last year of AI research; the

Explore this link on the map →

related reading