RL Explainer - Interactive Visualization of RL Algorithms for LLM Training
From REINFORCE to GRPO, DAPO, GSPO, VAPO, and beyond -- understand how each algorithm works, what changed in each formula, and why it matters for training reasoning models.
RL Explainer Interactive Visualization of Reinforcement Learning Algorithms for LLM Training From REINFORCE to GRPO, DAPO, GSPO, VAPO, and beyond -- understand how each algorithm works, what changed in each formula, and why it matters for training reasoning models. REINFORCEPPORLHFDPOGRPODAPOGSPOREINFORCE++VAPOGMPOGFPO Algorithm Evolution Timeline Click any algorithm to explore its details Foundation Critic-Free Preference-Based Advanced Training Pipeline See how data flows through the RL training loop Update weights & repeat GRPO Specific Components Reference ModelGroup of G…
saved by
related reading
- Lightweight Guide to understanding GRPO and RL principles - Musings of Muraligitlostmurali.com
- State of RL for reasoning LLMs | A. Weersaweers.de
- A vision researcher’s guide to some RL stuff: PPO & GRPO - Yuge (Jimmy) Shiyugeten.github.io
- GRPO++: Tricks for Making RL Actually Workcameronrwolfe.substack.com
- Why GRPO is Important and How it Worksghost.oxen.ai
- RLHF Bookrlhfbook.com
- A Guide to Reinforcement Learning Post-Training for LLMs: PPO, DPO, GRPO, and Beyondhuggingface.co
- DeepSeek-R1arxiv.org
- PPO for LLMs: A Guide for Normal Peoplecameronrwolfe.substack.com
- RLHF | John Lambertjohnwlambert.github.io
- RLHF & Post-Training Course by Nathan Lambertrlhfbook.com
- From REINFORCE to Dr. GRPOlancelqf.github.io