Self-Distillation Enables Continual Learning
This is experimental HTML to improve accessibility. We invite you to report rendering errors. Use Alt+Y to toggle on accessible reporting links and Alt+Shift+Y to toggle off. Learn more about this project and help improve conversions. Continual learning, enabling models to acquire new skills and knowledge without degrading existing capabilities, remains a fundamental challenge for foundation models. While on-policy reinforcement learning can reduce forgetting, it requires explicit reward functions that are often unavailable. Learning from expert demonstrations, the primary alternative, is dominated by supervised fine-tuning (SFT), which is inherently off-policy. We introduce Self-Distillation Fine-Tuning (SDFT), a simple method that enables on-policy learning directly from demonstrations. SDFT leverages in-context learning by using a demonstration-conditioned model as its own teacher, generating on-policy training signals that preserve prior capabilities while acquiring new skills. Ac
Self-Distillation Enables Continual Learning Idan Shenfeld 1 2 Mehul Damani 1 Jonas Hübotter 3 Pulkit Agrawal 1 2 1 MIT 2 Improbable AI Lab 3 ETH Zurich Correspondence to idanshen@mit.edu . Abstract Continual learning, enabling models to acquire new skills and knowledge without degrading existing capabilities, remains a fundamental challenge for foundation models. While on-policy reinforcement learning can reduce forgetting, it requires explicit reward functions that are often unavailable. Learning from expert demonstrations, the primary alternative, is dominated by supervised fine-tuning (SFT
Explore this link on the map →related reading
- On-Policy Distillation - Thinking Machines Labthinkingmachines.ai
- SFT, RL, and On-Policy Distillation Through a Distributional Lens | whnrehiew.github.io
- Why We Need Continual Learning | Andreessen Horowitza16z.com
- [2604.00626] A Survey of On-Policy Distillation for Large Language Modelsarxiv.org
- The Continual Learning Problemjessylin.com
- Bitter Lessons from Distillation Robustifies Unlearningbrucewlee.com
- [2601.18734] Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Modelsarxiv.org
- Pedagogical RL: Teaching Models to Teach Themselves from Privileged Information - Noah Ziemsnoahziems.com
- [2605.10889] Unmasking On-Policy Distillation: Where It Helps, Where It Hurts, and Whyarxiv.org
- will brown on X: "On SFT, RL, and on-policy distillation" / Xx.com
- augustus odena on X: "I have a bunch of thoughts about continual learning and nothing to do with them (I'm working on something else) so I figured I'd just turn them into a post: First: I think people use "continual learning" to point at a cluster of issues that are related but distinct. I'll list" / Xx.com
- What are the real problems of continual learning?infinitefaculty.substack.com