flâneur — a map of the web's best reading

From REINFORCE to Dr. GRPO

lancelqf.github.io · 3,340 words · saved by 1 readers

A Unified Perspective on LLM Post-Training

From REINFORCE to Dr. GRPO From REINFORCE to Dr. GRPO A Unified Perspective on LLM Post-Training March 27, 2025 Article Self Research Recently, many reinforcement learning (RL) algorithms have been applied to improve the post-training of large language models (LLMs). In this article, we aim to provide a unified perspective on the objectives of these RL algorithms, exploring how they relate to each other through the Policy Gradient Theorem [1] — the fundamental theorem of policy gradient methods. Background Let $\Delta(X)$ be the space of all probability distributions supported over the set $X$

Explore this link on the map →

saved by

related reading