Circumventing interpretability: How to defeat mind-readers — LessWrong
TL;DR: Unaligned AI will have a convergent instrumental incentive to make its thoughts difficult for us to interpret. In this article, I discuss many ways that a capable AI might circumvent scalable interpretability methods and suggest a framework for thinking about these risks. I categorize potential interpretability circumvention methods on three levels: Acknowledgements: I’m grateful to David Lindner, Evan R. Murphy, Alex Lintz, Sid Black, Kyle McDonnell, Laria Reynolds, Adam Shimi, and Daniel Braun whose comments greatly improved earlier drafts of this article. The article’s weaknesses are mine, but many of its strengths are due to their contributions. Additionally, this article benefited from the prior work of many authors, but especially: Evan Hubinger, Peter Barnett, Adam Shimi, Neel Nanda, Evan R. Murphy, Eliezer Yudkowsky, Chris Olah. I collected several of the potential circumvention methods from their work. This work was carried out while at Conjecture. There’s been a lot
x Circumventing interpretability: How to defeat mind-readers — LessWrong Conjecture (org) Interpretability (ML & AI) Security Mindset Instrumental convergence AI Frontpage 119 Circumventing interpretability: How to defeat mind-readers by Lee Sharkey 14th Jul 2022 AI Alignment Forum 39 min read 15 119 Ω 49 (Post now available as a pdf: https://arxiv.org/abs/2212.11415 ) TL;DR: Unaligned AI will have a convergent instrumental incentive to make its thoughts difficult for us to interpret. In this article, I discuss many ways that a capable AI might circumvent scalable interpretability methods and
Explore this link on the map →related reading
- Dario Amodei — The Urgency of Interpretabilitydarioamodei.com
- Neel Nanda on Mechanistic Interpretability: Progress, Limits, and Paths to Safer AI — EA Forumforum.effectivealtruism.org
- A Pragmatic Vision for Interpretability — AI Alignment Forumalignmentforum.org
- Tracing the thoughts of a large language model \ Anthropicanthropic.com
- Transformer Circuits Threadtransformer-circuits.pub
- An Ambitious Vision for Interpretability — AI Alignment Forumalignmentforum.org
- How Can Interpretability Researchers Help AGI Go Well? — AI Alignment Forumalignmentforum.org
- Against Almost Every Theory of Impact of Interpretability — AI Alignment Forumalignmentforum.org
- On Optimism for Interpretabilitygoodfire.ai
- Research Areas in Interpretability (The Alignment Project by UK AISI) — AI Alignment Forumalignmentforum.org
- Interpretability Will Not Reliably Find Deceptive AI — LessWronglesswrong.com
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org