flâneur — a map of the web's best reading

PPO for LLMs: A Guide for Normal People

cameronrwolfe.substack.com · 12,214 words · saved by 1 readers

Understanding the complex RL algorithm that gave us modern LLMs…

PPO for LLMs: A Guide for Normal People Understanding the complex RL algorithm that gave us modern LLMs… Cameron R. Wolfe, Ph.D. Oct 27, 2025 175 12 14 Share (from [4, 5, 8]) Over the last several years, reinforcement learning (RL) has been one of the most impactful areas of research for large language models (LLMs). Early research used RL to align LLMs to human preferences, and this initial work on applying RL to LLMs relied almost exclusively on Proximal Policy Optimization (PPO) [1]. This choice led PPO to become the default RL algorithm in LLM post-training for years— this is a long reign

Explore this link on the map →

saved by

related reading