flâneur — a map of the web's best reading

Richard Sutton – Father of RL thinks LLMs are a dead end

dwarkesh.com · 10,356 words · saved by 1 readers

Richard Sutton is the father of reinforcement learning, winner of the 2024 Turing Award, and author of The Bitter Lesson. And he thinks LLMs are a dead end. After interviewing him, my steel man of Richard’s position is this: LLMs aren’t capable of learning on-the-job, so no matter how much we scale, we’ll need some new architecture to enable continual learning. And once we have it, we won’t need a special training phase — the agent will just learn on-the-fly, like all humans, and indeed, like all animals. This new paradigm will render our current approach with LLMs obsolete. In our interview, I did my best to represent the view that LLMs might function as the foundation on which experiential learning can happen… Some sparks flew. A big thanks to the Alberta Machine Intelligence Institute for inviting me up to Edmonton and for letting me use their studio and equipment. Enjoy! Watch on YouTube; listen on Apple Podcasts or Spotify. Labelbox makes it possible to train AI agents in hyperrea

Playback speed × Share post Share post at current time Share from 0:00 0:00 / Generate transcript A transcript unlocks clips, previews, and editing. 130 34 29 Richard Sutton – Father of RL thinks LLMs are a dead end LLMs aren’t Bitter-Lesson-pilled Dwarkesh Patel Sep 26, 2025 130 34 29 Share Transcript Richard Sutton is the father of reinforcement learning, winner of the 2024 Turing Award, and author of The Bitter Lesson. And he thinks LLMs are a dead end. After interviewing him, my steel man of Richard’s position is this: LLMs aren’t capable of learning on-the-job, so no matter how much we sc

Explore this link on the map →

related reading