SFT, RL, and On-Policy Distillation Through a Distributional Lens | wh
I have been thinking about post-training methods in terms of distributions. A language model is a distribution over sequences. When we post-train it and attempt to teach it a task, we are reshaping this distribution. Different post-training methods differ in how they reshape this distribution, what they treat as the target and how directly they define this target.
I have been thinking about post-training methods in terms of distributions. A language model is a distribution over sequences. When we post-train it and attempt to teach it a task, we are reshaping this distribution. Different post-training methods differ in how they reshape this distribution, what they treat as the target and how directly they define this target. This is neither a very precise statement nor is it meant to be fully rigorous. I just find it to be a useful mental model, but I think it explains a lot of the qualitative differences between SFT, RL, and On-Policy Distillation. This
Explore this link on the map →saved by
related reading
- will brown on X: "On SFT, RL, and on-policy distillation" / Xx.com
- On-Policy Distillation - Thinking Machines Labthinkingmachines.ai
- ar0cket1 on X: "Solving OPSD (basically)" / Xx.com
- [2604.00626] A Survey of On-Policy Distillation for Large Language Modelsarxiv.org
- RL Scaling Laws for LLMs - by Cameron R. Wolfe, Ph.D.cameronrwolfe.substack.com
- [2604.13010] Lightning OPD: Efficient Post-Training for Large Reasoning Models with Offline On-Policy Distillationarxiv.org
- State of RL for reasoning LLMs | A. Weersaweers.de
- Pedagogical RL: Teaching Models to Teach Themselves from Privileged Information - Noah Ziemsnoahziems.com
- A vision researcher’s guide to some RL stuff: PPO & GRPO - Yuge (Jimmy) Shiyugeten.github.io
- [2605.10889] Unmasking On-Policy Distillation: Where It Helps, Where It Hurts, and Whyarxiv.org
- Self-Distillation Enables Continual Learningarxiv.org
- RLHF & Post-Training Course by Nathan Lambertrlhfbook.com