✳flâneur — a map of the web's best reading
From REINFORCE to Dr. GRPO
lancelqf.github.io · 3,340 words · saved by 1 readers
A Unified Perspective on LLM Post-Training
From REINFORCE to Dr. GRPO From REINFORCE to Dr. GRPO A Unified Perspective on LLM Post-Training March 27, 2025 Article Self Research Recently, many reinforcement learning (RL) algorithms have been applied to improve the post-training of large language models (LLMs). In this article, we aim to provide a unified perspective on the objectives of these RL algorithms, exploring how they relate to each other through the Policy Gradient Theorem [1] — the fundamental theorem of policy gradient methods. Background Let $\Delta(X)$ be the space of all probability distributions supported over the set $X$
Explore this link on the map →saved by
related reading
- A vision researcher’s guide to some RL stuff: PPO & GRPO - Yuge (Jimmy) Shiyugeten.github.io
- State of RL for reasoning LLMs | A. Weersaweers.de
- RLHF Bookrlhfbook.com
- A Guide to Reinforcement Learning Post-Training for LLMs: PPO, DPO, GRPO, and Beyondhuggingface.co
- PPO for LLMs: A Guide for Normal Peoplecameronrwolfe.substack.com
- Part 3: Intro to Policy Optimization - Spinning Up documentationspinningup.openai.com
- [2305.18290] Direct Preference Optimization: Your Language Model is Secretly a Reward Modelarxiv.org
- On-Policy Distillation - Thinking Machines Labthinkingmachines.ai
- RL Scaling Laws for LLMs - by Cameron R. Wolfe, Ph.D.cameronrwolfe.substack.com
- RLHF & Post-Training Course by Nathan Lambertrlhfbook.com
- Understanding R1-Zero-Like Training: A Critical Perspectivearxiv.org
- Lightweight Guide to understanding GRPO and RL principles - Musings of Muraligitlostmurali.com