✳flâneur — a map of the web's best reading
State of RL for reasoning LLMs | A. Weers
aweers.de · 5,052 words · saved by 6 readers
PhD student
State of RL for reasoning LLMs ¶ Reinforcement learning has been one of the most consequential additions to the LLM post-training stack. It was the key ingredient that transformed GPT-3 into InstructGPT [1] , and it has since become central to the current wave of reasoning improvements [2] [3] . The first generation of RL for LLMs was dominated by PPO [4] , a method developed for more conventional RL settings such as Atari games and robotics, but adapted very successfully to RLHF. The second generation, driven by the goal of improving reasoning capabilities, brought another round of algor
Explore this link on the map →saved by
related reading
- A vision researcher’s guide to some RL stuff: PPO & GRPO - Yuge (Jimmy) Shiyugeten.github.io
- RL Scaling Laws for LLMs - by Cameron R. Wolfe, Ph.D.cameronrwolfe.substack.com
- PPO for LLMs: A Guide for Normal Peoplecameronrwolfe.substack.com
- Xiuyu Li on X: "RL Interview Questions 2026" / Xx.com
- The State of Reinforcement Learning for LLM Reasoningmagazine.sebastianraschka.com
- DeepSeek-R1arxiv.org
- From REINFORCE to Dr. GRPOlancelqf.github.io
- A Guide to Reinforcement Learning Post-Training for LLMs: PPO, DPO, GRPO, and Beyondhuggingface.co
- The State of Reinforcement Learning for LLM Reasoningsebastianraschka.com
- RLHF Bookrlhfbook.com
- Understanding R1-Zero-Like Training: A Critical Perspectivearxiv.org
- RLHF: Reinforcement Learning from Human Feedbackhuyenchip.com