Just make the straw bigger
joanvelja.com · 8,049 words · saved by 1 readers
When life gives you one bit per lemon, how many of them will make for a lemonade?
Special thanks go to Tommaso Furlanello, Neev Parikh and @hallerite for reading earlier versions of this post and pushing me to write this. In LoRA without regret , Schulman et al. informally claim that supervised learning can deliver O ( #tokens ) O(\text{\#tokens}) O ( #tokens ) bits of supervision per episode, whereas policy-gradient RL gets only O ( 1 ) O(1) O ( 1 ) bits per episode because the advantage signal is effectively a single scalar per episode. Karpathy, too, at Dwarkesh’s podcast, compared RL to “ sucking supervision through a straw.” The informal claim here is that low-bandwidt
saved by
related reading
- LoRA Without Regret - Thinking Machines Labthinkingmachines.ai
- How can LLM RL Work Despite Information-Theoretic Inefficiencyberen.io
- State of RL for reasoning LLMs | A. Weersaweers.de
- RL is even more information inefficient than you thoughtdwarkesh.com
- The Extreme Inefficiency of RL for Frontier Models - Toby Ordtobyord.com
- RL is even more information inefficient than you thoughtsubstack.com
- A vision researcher’s guide to some RL stuff: PPO & GRPO - Yuge (Jimmy) Shiyugeten.github.io
- Speeding up RL with high-leverage samples | Applied Computeappliedcompute.com
- You Only Need Minimal RLVR Training: Extrapolating LLMs via Rank-1 Trajectoriesarxiv.org
- Interactive Visualization of RL Algorithms for LLM Trainingzcy233035.github.io
- Contra Dwarkesh on RL sample-efficiency via information theorynewsletter.danielpaleka.com
- Progressive Point Matchingprestonfu.com