flâneur — a map of the web's best reading

Distillation Robustifies Unlearning — LessWrong

lesswrong.com · 8,208 words · saved by 1 readers

Current “unlearning” methods only suppress capabilities instead of truly unlearning the capabilities. But if you distill an unlearned model into a randomly initialized model, the resulting network is actually robust to relearning. We show why this works, how well it works, and how to trade off compute for robustness. Produced as part of the ML Alignment & Theory Scholars Program in the winter 2024–25 cohort of the shard theory stream. Read our paper on ArXiv and enjoy an interactive demo. Maybe some future AI has long-term goals and humanity is in its way. Maybe future open-weight AIs have tons of bioterror expertise. If a system has dangerous knowledge, that system becomes more dangerous, either in the wrong hands or in the AI’s own “hands.” By making it harder to get AIs to share or use dangerous knowledge, we decrease (but do not eliminate) catastrophic risk. Misuse risk Robust unlearning prevents finetuning attacks from easily retraining a model to share or use the unlearned skill

x Distillation Robustifies Unlearning — LessWrong Machine Unlearning MATS Program Language Models (LLMs) AI Frontpage 2025 Top Fifty: 15 % 239 Distillation Robustifies Unlearning by Bruce W. Lee , Addie Foote , alexinf , leni , Jacob G-W , Harish Kamath , Bryce Woodworth , cloud , TurnTrout 13th Jun 2025 AI Alignment Forum Linkpost for arxiv.org 9 min read 43 239 Ω 88 Current “unlearning” methods only suppress capabilities instead of truly unlearning the capabilities. But if you distill an unlearned model into a randomly initialized model, the resulting network is actually robust to relearning

Explore this link on the map →

related reading