✳flâneur — a map of the web's best reading
Deep Forgetting & Unlearning for Safely-Scoped LLMs — AI Alignment Forum
alignmentforum.org · 4,771 words · saved by 1 readers
Thanks to Phillip Christoffersen, Adam Gleave, Anjali Gopal, Soroush Pour, and Fabien Roger for useful discussions and feedback. …
x Deep Forgetting & Unlearning for Safely-Scoped LLMs — AI Alignment Forum Machine Unlearning Adversarial Examples (AI) Adversarial Training Algorithms Interpretability (ML & AI) Open Problems Research Agendas AI Frontpage 57 Deep Forgetting & Unlearning for Safely-Scoped LLMs by scasper 5th Dec 2023 15 min read 30 57 Thanks to Phillip Christoffersen, Adam Gleave, Anjali Gopal, Soroush Pour, and Fabien Roger for useful discussions and feedback. TL;DR This post overviews a research agenda for avoiding unwanted latent capabilities in LLMs. It argues that "deep" forgetting and unlearning may be i
Explore this link on the map →related reading
- AI in 2025: gestalt — LessWronglesswrong.com
- Bitter Lessons from Distillation Robustifies Unlearningbrucewlee.com
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilitiesarxiv.org
- Modifying LLM Beliefs with Synthetic Document Finetuningalignment.anthropic.com
- [2401.05566] Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Trainingarxiv.org
- [2604.22082] Removing Sandbagging in LLMs by Training with Weak Supervisionarxiv.org
- [2604.16812] Introspection Adapters: Training LLMs to Report Their Learned Behaviorsarxiv.org
- Refusal in LLMs is mediated by a single direction — LessWronglesswrong.com
- Introspection Adapters: Training LLMs to Report Their Learned Behaviorsarxiv.org
- Should We Train Against (CoT) Monitors? — LessWronglesswrong.com
- Distillation Robustifies Unlearning — LessWronglesswrong.com