Bitter Lessons from Distillation Robustifies Unlearning
brucewlee.com · 2,067 words · saved by 4 readers
Bruce W. Lee. Machine Unlearning. Distillation. AI Safety.
Introduction My collaborators and I wrote a paper titled "Distillation Robustifies Unlearning" a few months ago. I'm writing this post to communicate what I think our paper actually says about the problem of unlearning. This is more of a personal account than a group statement. This post is also somewhat intuition-heavy because the goal is to describe the worldview that emerged from working on a paper. Additionally, I'll argue that distillation is an excellent opportunity for a safety intervention that also happens to align with economic incentives. Therefore, I think any future effort to unde
saved by
related reading
- Distillation Robustifies Unlearning — LessWronglesswrong.com
- On-Policy Distillation - Thinking Machines Labthinkingmachines.ai
- RL's Razor: Why Online Reinforcement Learning Forgets Lessarxiv.org
- Subliminal Learning: Language Models Transmit Behavioral Traits via Hidden Signals in Dataalignment.anthropic.com
- AI safety techniques leveraging distillation — LessWronglesswrong.com
- Self-Distillation Enables Continual Learningarxiv.org
- The distillation double bind: Distilling misaligned models either transfers misalignment or it doesn'tblog.redwoodresearch.org
- Incriminating misaligned AI models via distillation — LessWronglesswrong.com
- Self-Distillation Enables Continual Learningarxiv.org
- Self-Distillation Enables Continual Learningarxiv.org
- Deep Forgetting & Unlearning for Safely-Scoped LLMs — AI Alignment Forumalignmentforum.org
- Do your capabilities homework — LessWronglesswrong.com