How can LLM RL Work Despite Information-Theoretic Inefficiency
Epistemic Status: Obviously speculative and maybe obvious. The success of RL in LLMs has been puzzling me for a while. People have developed various information-theoretic style arguments by which they argue that RL is extremely informationally inefficient compared to pretraining, that it can only impart a tiny amount of bits,...
Epistemic Status : Obviously speculative and maybe obvious. The success of RL in LLMs has been puzzling me for a while. People have developed various information-theoretic style arguments by which they argue that RL is extremely informationally inefficient compared to pretraining , that it can only impart a tiny amount of bits, and that it can only bring out behaviours that are already in the base models etc. The basic intuition here is extremely obvious and essentially falls immediately out of the formulation of the two methods. Pretraining (and SFT etc) compute a loss on every token so every
saved by
- Ratan Kaliani
- Uzay Girit
- Vincent Cheng
- Dhruv Sheth
- Andrew Cai
- Emil Ryd
- Brady Gho
- Ishaan Panigrahi
- Al-Ekram Elahee Hridoy
- Chris Shi
- Agape Keleta
- Nathan Chen
related reading
- The Extreme Inefficiency of RL for Frontier Models - Toby Ordtobyord.com
- State of RL for reasoning LLMs | A. Weersaweers.de
- RL Scaling Laws for LLMs - by Cameron R. Wolfe, Ph.D.cameronrwolfe.substack.com
- RL is even more information inefficient than you thoughtdwarkesh.com
- RLHF: Reinforcement Learning from Human Feedbackhuyenchip.com
- RLHF & Post-Training Course by Nathan Lambertrlhfbook.com
- Why Does Reinforcement Learning Generalize? A Feature-Level Mechanistic Study of Post-Training in Large Language Modelsarxiv.org
- The upcoming GPT-3 moment for RL | Mechanize, Inc.mechanize.work
- Getting Caught Up to Modern LLM Research | Samarth Goeldev.samarthgoel.com
- GenAI Handbookgenai-handbook.github.io
- Andrej Karpathy — AGI is still a decade awaydwarkesh.com
- IsoCompute Playbook: Optimally Scaling Sampling Compute for RL Training of LLMscompute-optimal-rl-llm-scaling.github.io