will brown on X: "On SFT, RL, and on-policy distillation" / X
To view keyboard shortcuts, press question mark View keyboard shortcuts Article See new posts Conversation Prime Intellect reposted will brown @willccbb On SFT, RL, and on-policy distillation 13 31 314 22K Why the SFT-RL pipeline works, where on-policy distillation fits, and how self-distillation goes wrong. Authors: Will Brown & Claude Opus 4.7 April 30, 2026 [Editor's Note: Arguments are mine, writing is Claude's. This is partially an experiment in trying to get Claude to help speed-write and structure technical research blogs, drafted initially as artifacts and refined via "debate". I have too many blog ideas that I never get around to writing up, but the models finally feel good enough to help out with this (hopefully -- let me know what you think).] §1 — The standard pipeline and the compounding argument Most post-training pipelines are some version of "SFT first, then RL" — pre-train, supervised-finetune to get a baseline, then run RL once SFT data dries up or stops moving the
@willccbb: On SFT, RL, and on-policy distillation Why the SFT-RL pipeline works, where on-policy distillation fits, and how self-distillation goes wrong. Authors: Will Brown & Claude Opus 4.7 April 30, 2026 [Editor's Note: Arguments are mine, writing is Claude's. This is partially an experiment in trying to get Claude to help speed-write and structure technical research blogs, drafted initially as artifacts and refined via "debate". I have too many blog ideas that I never get around to writing up, but the models finally feel good enough to help out with this (hopefully -- let me know what y
Explore this link on the map →saved by
related reading
- SFT, RL, and On-Policy Distillation Through a Distributional Lens | whnrehiew.github.io
- On-Policy Distillation - Thinking Machines Labthinkingmachines.ai
- RL Scaling Laws for LLMs - by Cameron R. Wolfe, Ph.D.cameronrwolfe.substack.com
- Josh Engels on X: "New GDM interp research: SFT is a big deal for safety relevant behaviors. We recently investigated root causes for some of Gemini’s behaviors. We were surprised to find that many behaviors actually came from the initial supervised finetuning stage, not later stages like RL! 🧵 https://t.co/mLg87XuXK5" / Xx.com
- What I learned this week - Can distillation be stopped, Mythos and the cybersecurity equilibrium, Pipeline RLdwarkesh.com
- State of RL for reasoning LLMs | A. Weersaweers.de
- Is Frontier Asynchronous RL Solved? — Luke J. Huangluk-huang.github.io
- Pedagogical RL: Teaching Models to Teach Themselves from Privileged Information - Noah Ziemsnoahziems.com
- Self-Distillation Enables Continual Learningarxiv.org
- [2605.10889] Unmasking On-Policy Distillation: Where It Helps, Where It Hurts, and Whyarxiv.org
- [2604.00626] A Survey of On-Policy Distillation for Large Language Modelsarxiv.org
- [2601.18734] Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Modelsarxiv.org