Self-Distillation Enables Continual Learning
This is experimental HTML to improve accessibility. We invite you to report rendering errors. Use Alt+Y to toggle on accessible reporting links and Alt+Shift+Y to toggle off. Learn more about this project and help improve conversions. Continual learning, enabling models to acquire new skills and knowledge without degrading existing capabilities, remains a fundamental challenge for foundation models. While on-policy reinforcement learning can reduce forgetting, it requires explicit reward functions that are often unavailable. Learning from expert demonstrations, the primary alternative, is dominated by supervised fine-tuning (SFT), which is inherently off-policy. We introduce Self-Distillation Fine-Tuning (SDFT), a simple method that enables on-policy learning directly from demonstrations. SDFT leverages in-context learning by using a demonstration-conditioned model as its own teacher, generating on-policy training signals that preserve prior capabilities while acquiring new skills. Ac
Mehul Damani Jonas Hübotter Affiliation: MIT Improbable AI Lab ETH Zurich Pulkit Agrawal Abstract Continual learning, enabling models to acquire new skills and knowledge without degrading existing capabilities, remains a fundamental challenge for foundation models. While on-policy reinforcement learning can reduce forgetting, it requires explicit reward functions that are often unavailable. Learning from expert demonstrations, the primary alternative, is dominated by supervised fine-tuning (SFT), which is inherently off-policy. We introduce Self-Distillation Fine-Tuning (SDFT), a simple…
saved by
related reading
- Denser $\neq$ Better: Limits of On-Policy Self-Distillation for Continual Post-Trainingarxiv.org
- Self-Distillation Enables Continual Learningarxiv.org
- Reinforcement Learning via Self-Distillationarxiv.org
- Self-Distillation Enables Continual Learningarxiv.org
- Why Does Reinforcement Learning Generalize? A Feature-Level Mechanistic Study of Post-Training in Large Language Modelsarxiv.org
- On-Policy Distillation - Thinking Machines Labthinkingmachines.ai
- SFT, RL, and On-Policy Distillation Through a Distributional Lens | whnrehiew.github.io
- [2602.12275] On-Policy Context Distillation for Language Modelsarxiv.org
- Why We Need Continual Learning | Andreessen Horowitza16z.com
- RL's Razor: Why Online Reinforcement Learning Forgets Lessarxiv.org
- [2601.18734] Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Modelsarxiv.org
- The Continual Learning Problemjessylin.com