Distillation Robustifies Unlearning — LessWrong
Current “unlearning” methods only suppress capabilities instead of truly unlearning the capabilities. But if you distill an unlearned model into a randomly initialized model, the resulting network is actually robust to relearning. We show why this works, how well it works, and how to trade off compute for robustness. Produced as part of the ML Alignment & Theory Scholars Program in the winter 2024–25 cohort of the shard theory stream. Read our paper on ArXiv and enjoy an interactive demo. Maybe some future AI has long-term goals and humanity is in its way. Maybe future open-weight AIs have tons of bioterror expertise. If a system has dangerous knowledge, that system becomes more dangerous, either in the wrong hands or in the AI’s own “hands.” By making it harder to get AIs to share or use dangerous knowledge, we decrease (but do not eliminate) catastrophic risk. Misuse risk Robust unlearning prevents finetuning attacks from easily retraining a model to share or use the unlearned skill
x Distillation Robustifies Unlearning — LessWrong Machine Unlearning MATS Program Language Models (LLMs) AI Frontpage 2025 Top Fifty: 15 % 239 Distillation Robustifies Unlearning by Bruce W. Lee , Addie Foote , alexinf , leni , Jacob G-W , Harish Kamath , Bryce Woodworth , cloud , TurnTrout 13th Jun 2025 AI Alignment Forum Linkpost for arxiv.org 9 min read 43 239 Ω 88 Current “unlearning” methods only suppress capabilities instead of truly unlearning the capabilities. But if you distill an unlearned model into a randomly initialized model, the resulting network is actually robust to relearning
Explore this link on the map →related reading
- Bitter Lessons from Distillation Robustifies Unlearningbrucewlee.com
- On-Policy Distillation - Thinking Machines Labthinkingmachines.ai
- AI safety techniques leveraging distillation — LessWronglesswrong.com
- Subliminal Learning: Language Models Transmit Behavioral Traits via Hidden Signals in Dataalignment.anthropic.com
- [1511.03643] Unifying distillation and privileged informationarxiv.org
- Deep Forgetting & Unlearning for Safely-Scoped LLMs — AI Alignment Forumalignmentforum.org
- What I learned this week - Can distillation be stopped, Mythos and the cybersecurity equilibrium, Pipeline RLdwarkesh.com
- Self-Distillation Enables Continual Learningarxiv.org
- SFT, RL, and On-Policy Distillation Through a Distributional Lens | whnrehiew.github.io
- [2605.10889] Unmasking On-Policy Distillation: Where It Helps, Where It Hurts, and Whyarxiv.org
- Incriminating misaligned AI models via distillation — LessWronglesswrong.com
- Distillation Walkthroughvladfeinberg.com