✳flâneur — a map of the web's best reading
PPO for LLMs: A Guide for Normal People
cameronrwolfe.substack.com · 12,214 words · saved by 1 readers
Understanding the complex RL algorithm that gave us modern LLMs…
PPO for LLMs: A Guide for Normal People Understanding the complex RL algorithm that gave us modern LLMs… Cameron R. Wolfe, Ph.D. Oct 27, 2025 175 12 14 Share (from [4, 5, 8]) Over the last several years, reinforcement learning (RL) has been one of the most impactful areas of research for large language models (LLMs). Early research used RL to align LLMs to human preferences, and this initial work on applying RL to LLMs relied almost exclusively on Proximal Policy Optimization (PPO) [1]. This choice led PPO to become the default RL algorithm in LLM post-training for years— this is a long reign
Explore this link on the map →saved by
related reading
- State of RL for reasoning LLMs | A. Weersaweers.de
- A vision researcher’s guide to some RL stuff: PPO & GRPO - Yuge (Jimmy) Shiyugeten.github.io
- [1707.06347] Proximal Policy Optimization Algorithmsarxiv.org
- The 37 Implementation Details of Proximal Policy Optimization · The ICLR Blog Trackiclr-blog-track.github.io
- A Guide to Reinforcement Learning Post-Training for LLMs: PPO, DPO, GRPO, and Beyondhuggingface.co
- RLHF Bookrlhfbook.com
- From REINFORCE to Dr. GRPOlancelqf.github.io
- Rethinking the Role of PPO in RLHF – The Berkeley Artificial Intelligence Research Blogbair.berkeley.edu
- RL Scaling Laws for LLMs - by Cameron R. Wolfe, Ph.D.cameronrwolfe.substack.com
- [2305.18290] Direct Preference Optimization: Your Language Model is Secretly a Reward Modelarxiv.org
- RLHF & Post-Training Course by Nathan Lambertrlhfbook.com
- On-Policy Distillation - Thinking Machines Labthinkingmachines.ai