flâneur — a map of the web's best reading

Dwarkesh Patel on X: "The most interesting part for me is where @karpathy describes why LLMs aren't able to learn like humans. As you would expect, he comes up with a wonderfully evocative phrase to describe RL: “sucking supervision bits through a straw.” A single end reward gets broadcast across https://t.co/lYonLgrukB" / X

x.com · 699 words · saved by 1 readers

To view keyboard shortcuts, press question mark View keyboard shortcuts Messages Messages Chat Beta Message requests 10+ new requests Emil Ryd @emilaryd · Oct 9 Shared a post Chana @ChanaMessinger · Sep 30 Probably not very interested in this but I'd want to hear more nonetheless. Pliny the Liberator 󠅫󠄼󠄿󠅆󠄵󠄐󠅀󠄼󠄹󠄾󠅉󠅭 @elder_plinius · Sep 24 Hey Pliny had a quick question. On average, when jailbreaking a new model, how many times do you have to try before you are successful? Rithvik @Rithvik43293318 · Sep 2 supp Luke Drago @luke_drago_ · Sep 2 You reacted to a post with kool123r5 @kool123r5 · Jul 23 You reacted with : Og (I have no idea what that is) risha :) @swankypizza123 · Jul 23 lol atreyi @atreyisxha · Jul 23 hi yuvika @yuuvika · Jul 23 yo Rob Wiblin @robertwiblin · Jul 16 I see, thanks for the response! 18 8th grade · May 8 Aneesh: ? Sonith @_sonith · Mar 31 You sent a link 2 Hackathon · Dec 26, 2023 KingL☆ reacted to @Arjunkh07’s message with : twttier not for dms 11 9t

Dwarkesh Patel @dwarkesh_sp The most interesting part for me is where @ karpathy describes why LLMs aren't able to learn like humans. As you would expect, he comes up with a wonderfully evocative phrase to describe RL: “sucking supervision bits through a straw.” A single end reward gets broadcast across every token in a successful trajectory, upweighting even wrong or irrelevant turns that lead to the right answer. > “Humans don't use reinforcement learning, as I've said before. I think they do something different. Reinforcement learning is a lot worse than the average person thinks. Reinforce

Explore this link on the map →

related reading