How can LLM RL Work Despite Information-Theoretic Inefficiency
Epistemic Status: Obviously speculative and maybe obvious. The success of RL in LLMs has been puzzling me for a while. People have developed various information-theoretic style arguments by which they argue that RL is extremely informationally inefficient compared to pretraining, that it can only impart a tiny amount of bits,...
Epistemic Status : Obviously speculative and maybe obvious. The success of RL in LLMs has been puzzling me for a while. People have developed various information-theoretic style arguments by which they argue that RL is extremely informationally inefficient compared to pretraining , that it can only impart a tiny amount of bits, and that it can only bring out behaviours that are already in the base models etc. The basic intuition here is extremely obvious and essentially falls immediately out of the formulation of the two methods. Pretraining (and SFT etc) compute a loss on every token so every
Explore this link on the map →related reading
- State of RL for reasoning LLMs | A. Weersaweers.de
- RL is even more information inefficient than you thoughtdwarkesh.com
- RLHF: Reinforcement Learning from Human Feedbackhuyenchip.com
- RL Scaling Laws for LLMs - by Cameron R. Wolfe, Ph.D.cameronrwolfe.substack.com
- RLHF & Post-Training Course by Nathan Lambertrlhfbook.com
- The Extreme Inefficiency of RL for Frontier Models - Toby Ordtobyord.com
- Andrej Karpathy — AGI is still a decade awaydwarkesh.com
- GenAI Handbookgenai-handbook.github.io
- A vision researcher’s guide to some RL stuff: PPO & GRPO - Yuge (Jimmy) Shiyugeten.github.io
- IsoCompute Playbook: Optimally Scaling Sampling Compute for RL Training of LLMscompute-optimal-rl-llm-scaling.github.io
- PPO for LLMs: A Guide for Normal Peoplecameronrwolfe.substack.com
- LLM Resourcesforrestbicker.com