[2601.19897] Self-Distillation Enables Continual Learning
Abstract:Continual learning, enabling models to acquire new skills and knowledge without degrading existing capabilities, remains a fundamental challenge for foundation models. While on-policy reinforcement learning can reduce forgetting, it requires explicit reward functions that are often unavailable. Learning from expert demonstrations, the primary alternative, is dominated by supervised fine-tuning (SFT), which is inherently off-policy. We introduce Self-Distillation Fine-Tuning (SDFT), a simple method that enables on-policy learning directly from demonstrations. SDFT leverages in-context learning by using a demonstration-conditioned model as its own teacher, generating on-policy training signals that preserve prior capabilities while acquiring new skills. Across skill learning and knowledge acquisition tasks, SDFT consistently outperforms SFT, achieving higher new-task accuracy while substantially reducing catastrophic forgetting. In sequential learning experiments, SDFT enables a single model to accumulate multiple skills over time without performance regression, establishing on-policy distillation as a practical path to continual learning from demonstrations.
View PDF HTML (experimental) Abstract:Continual learning, enabling models to acquire new skills and knowledge without degrading existing capabilities, remains a fundamental challenge for foundation models. While on-policy reinforcement learning can reduce forgetting, it requires explicit reward functions that are often unavailable. Learning from expert demonstrations, the primary alternative, is dominated by supervised fine-tuning (SFT), which is inherently off-policy. We introduce Self-Distillation Fine-Tuning (SDFT), a simple method that enables on-policy learning directly from…
saved by
related reading
- Self-Distillation Enables Continual Learningarxiv.org
- Self-Distillation Enables Continual Learningarxiv.org
- On-Policy Distillation - Thinking Machines Labthinkingmachines.ai
- Denser $\neq$ Better: Limits of On-Policy Self-Distillation for Continual Post-Trainingarxiv.org
- Introducing S1: In-Context Learning for Roboticsskild.ai
- The Continual Learning Problemjessylin.com
- RL's Razor: Why Online Reinforcement Learning Forgets Lessarxiv.org
- Why We Need Continual Learning | Andreessen Horowitza16z.com
- GEN-1.5: Embodied Foundation Models are One-Shot Learners - Generalist AIgeneralistai.com
- SFT, RL, and On-Policy Distillation Through a Distributional Lens | whnrehiew.github.io
- [2602.12275] On-Policy Context Distillation for Language Modelsarxiv.org
- Bitter Lessons from Distillation Robustifies Unlearningbrucewlee.com