PPO for LLMs: A Guide for Normal People
cameronrwolfe.substack.com · 12,214 words · saved by 1 readers
Understanding the complex RL algorithm that gave us modern LLMs…
PPO for LLMs: A Guide for Normal People Understanding the complex RL algorithm that gave us modern LLMs… Cameron R. Wolfe, Ph.D. Oct 27, 2025 175 12 14 Share (from [4, 5, 8]) Over the last several years, reinforcement learning (RL) has been one of the most impactful areas of research for large language models (LLMs). Early research used RL to align LLMs to human preferences, and this initial work on applying RL to LLMs relied almost exclusively on Proximal Policy Optimization (PPO) [1]. This choice led PPO to become the default RL algorithm in LLM post-training for years— this is a long reign
saved by
related reading
- State of RL for reasoning LLMs | A. Weersaweers.de
- GRPO++: Tricks for Making RL Actually Workcameronrwolfe.substack.com
- RLHF | John Lambertjohnwlambert.github.io
- Understanding Policy Gradients | John Lambertjohnwlambert.github.io
- [2602.19362] LLMs Can Learn to Reason Via Off-Policy RLarxiv.org
- A vision researcher’s guide to some RL stuff: PPO & GRPO - Yuge (Jimmy) Shiyugeten.github.io
- [1707.06347] Proximal Policy Optimization Algorithmsarxiv.org
- The 37 Implementation Details of Proximal Policy Optimization · The ICLR Blog Trackiclr-blog-track.github.io
- A Guide to Reinforcement Learning Post-Training for LLMs: PPO, DPO, GRPO, and Beyondhuggingface.co
- RLHF Bookrlhfbook.com
- From REINFORCE to Dr. GRPOlancelqf.github.io
- Harsh Bhatt (@harshbhatt7585) on Xx.com