Deep Forgetting & Unlearning for Safely-Scoped LLMs — AI Alignment Forum
alignmentforum.org · 4,771 words · saved by 1 readers
Thanks to Phillip Christoffersen, Adam Gleave, Anjali Gopal, Soroush Pour, and Fabien Roger for useful discussions and feedback. …
x Deep Forgetting & Unlearning for Safely-Scoped LLMs — AI Alignment Forum Machine Unlearning Adversarial Examples (AI) Adversarial Training Algorithms Interpretability (ML & AI) Open Problems Research Agendas AI Frontpage 57 Deep Forgetting & Unlearning for Safely-Scoped LLMs by scasper 5th Dec 2023 15 min read 30 57 Thanks to Phillip Christoffersen, Adam Gleave, Anjali Gopal, Soroush Pour, and Fabien Roger for useful discussions and feedback. TL;DR This post overviews a research agenda for avoiding unwanted latent capabilities in LLMs. It argues that "deep" forgetting and unlearning may be i
related reading
- AI in 2025: gestalt — LessWronglesswrong.com
- Bitter Lessons from Distillation Robustifies Unlearningbrucewlee.com
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- [2502.05209] Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilitiesarxiv.org
- Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilitiesarxiv.org
- [2512.05648] Beyond Data Filtering: Knowledge Localization for Capability Removal in LLMsarxiv.org
- [2401.05566] Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Trainingarxiv.org
- Refusal in LLMs is mediated by a single direction — LessWronglesswrong.com
- Introspection Adapters: Training LLMs to Report Their Learned Behaviorsarxiv.org
- [2604.16812] Introspection Adapters: Training LLMs to Report Their Learned Behaviorsarxiv.org
- Student Projects - CS 2881R AI Safetyboazbk.github.io
- [2512.11949] Neural Chameleons: Language Models Can Learn to Hide Their Thoughts from Unseen Activation Monitorsarxiv.org