flâneur — a map of the web's best reading

will brown on X: "On SFT, RL, and on-policy distillation" / X

x.com · 628 words · saved by 1 readers

To view keyboard shortcuts, press question mark View keyboard shortcuts Article See new posts Conversation Prime Intellect reposted will brown @willccbb On SFT, RL, and on-policy distillation 13 31 314 22K Why the SFT-RL pipeline works, where on-policy distillation fits, and how self-distillation goes wrong. Authors: Will Brown & Claude Opus 4.7 April 30, 2026 [Editor's Note: Arguments are mine, writing is Claude's. This is partially an experiment in trying to get Claude to help speed-write and structure technical research blogs, drafted initially as artifacts and refined via "debate". I have too many blog ideas that I never get around to writing up, but the models finally feel good enough to help out with this (hopefully -- let me know what you think).] §1 — The standard pipeline and the compounding argument Most post-training pipelines are some version of "SFT first, then RL" — pre-train, supervised-finetune to get a baseline, then run RL once SFT data dries up or stops moving the

@willccbb: On SFT, RL, and on-policy distillation Why the SFT-RL pipeline works, where on-policy distillation fits, and how self-distillation goes wrong. Authors: Will Brown & Claude Opus 4.7 April 30, 2026 [Editor's Note: Arguments are mine, writing is Claude's. This is partially an experiment in trying to get Claude to help speed-write and structure technical research blogs, drafted initially as artifacts and refined via "debate". I have too many blog ideas that I never get around to writing up, but the models finally feel good enough to help out with this (hopefully -- let me know what y

Explore this link on the map →

saved by

related reading