Dense, on-policy, or both?
baseten.co · 2,441 words · saved by 2 readers
Constitutional alignment as a testbed for comparing learning signals in SFT, RL, and everything in between
Introduction Post-training methods can be usefully organised along two axes: whether the model trains on its own generations (on-policy) or someone else's (off-policy), and whether the learning signal is dense (token-level) or sparse (sequence-level). Much of our recent work has tried to populate the space between off-policy SFT and on-policy RL. Iterative SFT (iSFT), which converts supervised data into gold-standard samples via a grader- and refiner-loop, moves SFT closer to the current policy by repairing model generations; RL optimizes directly on policy but with scalar rewards; and…
saved by
related reading
- [2601.18734] Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Modelsarxiv.org
- On-Policy Distillation - Thinking Machines Labthinkingmachines.ai
- SFT, RL, and On-Policy Distillation Through a Distributional Lens | whnrehiew.github.io
- will brown on X: "On SFT, RL, and on-policy distillation" / Xx.com
- Synthetic Persona Pretraining: Alignment from Token Zeromodelraising.ai
- On-Policy Distillation: Promise, Pitfalls, and Prospectslouieworth.github.io
- State of RL for reasoning LLMs | A. Weersaweers.de
- Self-Distillation Enables Continual Learningarxiv.org
- [2605.10889] Unmasking On-Policy Distillation: Where It Helps, Where It Hurts, and Whyarxiv.org
- Denser $\neq$ Better: Limits of On-Policy Self-Distillation for Continual Post-Trainingarxiv.org
- On the Geometry of On-Policy Distillationarxiv.org
- Pedagogical RL: Teaching Models to Teach Themselves from Privileged Information - Noah Ziemsnoahziems.com