flâneur — a map of the web's best reading

Deep Forgetting & Unlearning for Safely-Scoped LLMs — AI Alignment Forum

alignmentforum.org · 4,771 words · saved by 1 readers

Thanks to Phillip Christoffersen, Adam Gleave, Anjali Gopal, Soroush Pour, and Fabien Roger for useful discussions and feedback. …

x Deep Forgetting & Unlearning for Safely-Scoped LLMs — AI Alignment Forum Machine Unlearning Adversarial Examples (AI) Adversarial Training Algorithms Interpretability (ML & AI) Open Problems Research Agendas AI Frontpage 57 Deep Forgetting & Unlearning for Safely-Scoped LLMs by scasper 5th Dec 2023 15 min read 30 57 Thanks to Phillip Christoffersen, Adam Gleave, Anjali Gopal, Soroush Pour, and Fabien Roger for useful discussions and feedback. TL;DR This post overviews a research agenda for avoiding unwanted latent capabilities in LLMs. It argues that "deep" forgetting and unlearning may be i

Explore this link on the map →

related reading