flâneur

RL Explainer - Interactive Visualization of RL Algorithms for LLM Training

zcy233035.github.io · 1,079 words · saved by 1 readers

From REINFORCE to GRPO, DAPO, GSPO, VAPO, and beyond -- understand how each algorithm works, what changed in each formula, and why it matters for training reasoning models.

RL Explainer Interactive Visualization of Reinforcement Learning Algorithms for LLM Training From REINFORCE to GRPO, DAPO, GSPO, VAPO, and beyond -- understand how each algorithm works, what changed in each formula, and why it matters for training reasoning models. REINFORCEPPORLHFDPOGRPODAPOGSPOREINFORCE++VAPOGMPOGFPO Algorithm Evolution Timeline Click any algorithm to explore its details Foundation Critic-Free Preference-Based Advanced Training Pipeline See how data flows through the RL training loop Update weights & repeat GRPO Specific Components Reference ModelGroup of G…

saved by

related reading