Richard Sutton – Father of RL thinks LLMs are a dead end
Richard Sutton is the father of reinforcement learning, winner of the 2024 Turing Award, and author of The Bitter Lesson. And he thinks LLMs are a dead end. After interviewing him, my steel man of Richard’s position is this: LLMs aren’t capable of learning on-the-job, so no matter how much we scale, we’ll need some new architecture to enable continual learning. And once we have it, we won’t need a special training phase — the agent will just learn on-the-fly, like all humans, and indeed, like all animals. This new paradigm will render our current approach with LLMs obsolete. In our interview, I did my best to represent the view that LLMs might function as the foundation on which experiential learning can happen… Some sparks flew. A big thanks to the Alberta Machine Intelligence Institute for inviting me up to Edmonton and for letting me use their studio and equipment. Enjoy! Watch on YouTube; listen on Apple Podcasts or Spotify. Labelbox makes it possible to train AI agents in hyperrea
Playback speed × Share post Share post at current time Share from 0:00 0:00 / Generate transcript A transcript unlocks clips, previews, and editing. 130 34 29 Richard Sutton – Father of RL thinks LLMs are a dead end LLMs aren’t Bitter-Lesson-pilled Dwarkesh Patel Sep 26, 2025 130 34 29 Share Transcript Richard Sutton is the father of reinforcement learning, winner of the 2024 Turing Award, and author of The Bitter Lesson. And he thinks LLMs are a dead end. After interviewing him, my steel man of Richard’s position is this: LLMs aren’t capable of learning on-the-job, so no matter how much we sc
Explore this link on the map →related reading
- Andrej Karpathy — AGI is still a decade awaydwarkesh.com
- Animals vs Ghosts – karpathykarpathy.bearblog.dev
- The Era of Experience Paper.pdfstorage.googleapis.com
- 2025 LLM Year in Review – karpathykarpathy.bearblog.dev
- Andrej Karpathy on X: "Finally had a chance to listen through this pod with Sutton, which was interesting and amusing. As background, Sutton's "The Bitter Lesson" has become a bit of biblical text in frontier LLM circles. Researchers routinx.com
- RLHF: Reinforcement Learning from Human Feedbackhuyenchip.com
- GenAI Handbookgenai-handbook.github.io
- Just Ask for Generalization | Eric Jangevjang.com
- State of RL for reasoning LLMs | A. Weersaweers.de
- The man who taught AI to learn believes human-level intelligence is closer than you think | IBMibm.com
- Against LLM Reductionism — LessWronglesswrong.com
- Dwarkesh Patel on X: "The most interesting part for me is where @karpathy describes why LLMs aren't able to learn like humans. As you would expect, he comes up with a wonderfully evocative phrase to describe RL: “sucking supervision bits thx.com