✳flâneur — a map of the web's best reading
Bitter Lessons from Distillation Robustifies Unlearning
brucewlee.com · 2,067 words · saved by 4 readers
Bruce W. Lee. Machine Unlearning. Distillation. AI Safety.
Introduction My collaborators and I wrote a paper titled "Distillation Robustifies Unlearning" a few months ago. I'm writing this post to communicate what I think our paper actually says about the problem of unlearning. This is more of a personal account than a group statement. This post is also somewhat intuition-heavy because the goal is to describe the worldview that emerged from working on a paper. Additionally, I'll argue that distillation is an excellent opportunity for a safety intervention that also happens to align with economic incentives. Therefore, I think any future effort to unde
Explore this link on the map →saved by
related reading
- Distillation Robustifies Unlearning — LessWronglesswrong.com
- On-Policy Distillation - Thinking Machines Labthinkingmachines.ai
- Subliminal Learning: Language Models Transmit Behavioral Traits via Hidden Signals in Dataalignment.anthropic.com
- AI safety techniques leveraging distillation — LessWronglesswrong.com
- Self-Distillation Enables Continual Learningarxiv.org
- Deep Forgetting & Unlearning for Safely-Scoped LLMs — AI Alignment Forumalignmentforum.org
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- SFT, RL, and On-Policy Distillation Through a Distributional Lens | whnrehiew.github.io
- Incriminating misaligned AI models via distillation — LessWronglesswrong.com
- Selective modularity: a research agenda — LessWronglesswrong.com
- Why We Need Continual Learning | Andreessen Horowitza16z.com
- [2605.10889] Unmasking On-Policy Distillation: Where It Helps, Where It Hurts, and Whyarxiv.org