Dwarkesh Patel on X: "The most interesting part for me is where @karpathy describes why LLMs aren't able to learn like humans. As you would expect, he comes up with a wonderfully evocative phrase to describe RL: “sucking supervision bits through a straw.” A single end reward gets broadcast across https://t.co/lYonLgrukB" / X
To view keyboard shortcuts, press question mark View keyboard shortcuts Messages Messages Chat Beta Message requests 10+ new requests Emil Ryd @emilaryd · Oct 9 Shared a post Chana @ChanaMessinger · Sep 30 Probably not very interested in this but I'd want to hear more nonetheless. Pliny the Liberator 󠅫󠄼󠄿󠅆󠄵󠄐󠅀󠄼󠄹󠄾󠅉󠅭 @elder_plinius · Sep 24 Hey Pliny had a quick question. On average, when jailbreaking a new model, how many times do you have to try before you are successful? Rithvik @Rithvik43293318 · Sep 2 supp Luke Drago @luke_drago_ · Sep 2 You reacted to a post with kool123r5 @kool123r5 · Jul 23 You reacted with : Og (I have no idea what that is) risha :) @swankypizza123 · Jul 23 lol atreyi @atreyisxha · Jul 23 hi yuvika @yuuvika · Jul 23 yo Rob Wiblin @robertwiblin · Jul 16 I see, thanks for the response! 18 8th grade · May 8 Aneesh: ? Sonith @_sonith · Mar 31 You sent a link 2 Hackathon · Dec 26, 2023 KingL☆ reacted to @Arjunkh07’s message with : twttier not for dms 11 9t
Dwarkesh Patel @dwarkesh_sp The most interesting part for me is where @ karpathy describes why LLMs aren't able to learn like humans. As you would expect, he comes up with a wonderfully evocative phrase to describe RL: “sucking supervision bits through a straw.” A single end reward gets broadcast across every token in a successful trajectory, upweighting even wrong or irrelevant turns that lead to the right answer. > “Humans don't use reinforcement learning, as I've said before. I think they do something different. Reinforcement learning is a lot worse than the average person thinks. Reinforce
Explore this link on the map →related reading
- Andrej Karpathy — AGI is still a decade awaydwarkesh.com
- Animals vs Ghosts – karpathykarpathy.bearblog.dev
- RLHF: Reinforcement Learning from Human Feedbackhuyenchip.com
- LLM Daydreaming · Gwern.netgwern.net
- 2025 LLM Year in Review – karpathykarpathy.bearblog.dev
- RLHF & Post-Training Course by Nathan Lambertrlhfbook.com
- Guardian Angels: LLM Personalization for Productivity and Security · Gwern.netgwern.net
- GenAI Handbookgenai-handbook.github.io
- Language Models in Plato's Cave - by Sergey Levinesergeylevine.substack.com
- State of RL for reasoning LLMs | A. Weersaweers.de
- Richard Sutton – Father of RL thinks LLMs are a dead enddwarkesh.com
- Andrej Karpathy on X: "Finally had a chance to listen through this pod with Sutton, which was interesting and amusing. As background, Sutton's "The Bitter Lesson" has become a bit of biblical text in frontier LLM circles. Researchers routinx.com