State of RL for reasoning LLMs | A. Weers
aweers.de · 5,052 words · saved by 6 readers
PhD student
State of RL for reasoning LLMs ¶ Reinforcement learning has been one of the most consequential additions to the LLM post-training stack. It was the key ingredient that transformed GPT-3 into InstructGPT [1] , and it has since become central to the current wave of reasoning improvements [2] [3] . The first generation of RL for LLMs was dominated by PPO [4] , a method developed for more conventional RL settings such as Atari games and robotics, but adapted very successfully to RLHF. The second generation, driven by the goal of improving reasoning capabilities, brought another round of algor
saved by
related reading
- GRPO++: Tricks for Making RL Actually Workcameronrwolfe.substack.com
- A vision researcher’s guide to some RL stuff: PPO & GRPO - Yuge (Jimmy) Shiyugeten.github.io
- RL Scaling Laws for LLMs - by Cameron R. Wolfe, Ph.D.cameronrwolfe.substack.com
- PPO for LLMs: A Guide for Normal Peoplecameronrwolfe.substack.com
- How can LLM RL Work Despite Information-Theoretic Inefficiencyberen.io
- [2602.19362] LLMs Can Learn to Reason Via Off-Policy RLarxiv.org
- DeepSeek-R1arxiv.org
- Xiuyu Li on X: "RL Interview Questions 2026" / Xx.com
- RLHF | John Lambertjohnwlambert.github.io
- The State of Reinforcement Learning for LLM Reasoningmagazine.sebastianraschka.com
- RL ALGOk-a.in
- Progressive Point Matchingprestonfu.com