Against Almost Every Theory of Impact of Interpretability — AI Alignment Forum
Epistemic Status: I believe I am well-versed in this subject. I erred on the side of making claims that were too strong and allowing readers to disagree and start a discussion about precise points rather than trying to edge-case every statement. I also think that using memes is important because safety ideas are boring and anti-memetic. So let’s go! Many thanks to @scasper, @Sid Black , @Neel Nanda , @Fabien Roger , @Bogdan Ionut Cirstea, @WCargo, @Alexandre Variengien, @Jonathan Claybrough, @Edoardo Pona, @Andrea_Miotti, Diego Dorn, Angélina Gentaz, Clement Dumas, and Enzo Marsot for useful feedback and discussions. When I started this post, I began by critiquing the article A Long List of Theories of Impact for Interpretability, from Neel Nanda, but I later expanded the scope of my critique. Some ideas which are presented are not supported by anyone, but to explain the difficulties, I still need to 1. explain them and 2. criticize them. It gives an adversarial vibe to this post. I'm
x Against Almost Every Theory of Impact of Interpretability — AI Alignment Forum Best of LessWrong 2023 Interpretability (ML & AI) AI Frontpage 96 Against Almost Every Theory of Impact of Interpretability by Charbel-Raphaël 17th Aug 2023 31 min read 93 96 Epistemic Status: I believe I am well-versed in this subject. I erred on the side of making claims that were too strong and allowing readers to disagree and start a discussion about precise points rather than trying to edge-case every statement. I also think that using memes is important because safety ideas are boring and anti-memetic . So l
Explore this link on the map →related reading
- Against Almost Every Theory of Impact of Interpretability — LessWronglesswrong.com
- Dario Amodei — The Urgency of Interpretabilitydarioamodei.com
- A Pragmatic Vision for Interpretability — AI Alignment Forumalignmentforum.org
- Neel Nanda on Mechanistic Interpretability: Progress, Limits, and Paths to Safer AI — EA Forumforum.effectivealtruism.org
- How Can Interpretability Researchers Help AGI Go Well? — AI Alignment Forumalignmentforum.org
- Interpretability Will Not Reliably Find Deceptive AI — AI Alignment Forumalignmentforum.org
- The Hitchhiker's Guide to Actionable Interpretabilityactionable-interpretability-guide.github.io
- On Optimism for Interpretabilitygoodfire.ai
- An Ambitious Vision for Interpretability — AI Alignment Forumalignmentforum.org
- Interpretability Will Not Reliably Find Deceptive AI — LessWronglesswrong.com
- How To Become A Mechanistic Interpretability Researcher — AI Alignment Forumalignmentforum.org
- Research Areas in Interpretability (The Alignment Project by UK AISI) — AI Alignment Forumalignmentforum.org