flâneur — a map of the web's best reading

Circumventing interpretability: How to defeat mind-readers — LessWrong

lesswrong.com · 13,224 words · saved by 1 readers

TL;DR: Unaligned AI will have a convergent instrumental incentive to make its thoughts difficult for us to interpret. In this article, I discuss many ways that a capable AI might circumvent scalable interpretability methods and suggest a framework for thinking about these risks. I categorize potential interpretability circumvention methods on three levels: Acknowledgements: I’m grateful to David Lindner, Evan R. Murphy, Alex Lintz, Sid Black, Kyle McDonnell, Laria Reynolds, Adam Shimi, and Daniel Braun whose comments greatly improved earlier drafts of this article. The article’s weaknesses are mine, but many of its strengths are due to their contributions. Additionally, this article benefited from the prior work of many authors, but especially: Evan Hubinger, Peter Barnett, Adam Shimi, Neel Nanda, Evan R. Murphy, Eliezer Yudkowsky, Chris Olah. I collected several of the potential circumvention methods from their work. This work was carried out while at Conjecture. There’s been a lot

x Circumventing interpretability: How to defeat mind-readers — LessWrong Conjecture (org) Interpretability (ML & AI) Security Mindset Instrumental convergence AI Frontpage 119 Circumventing interpretability: How to defeat mind-readers by Lee Sharkey 14th Jul 2022 AI Alignment Forum 39 min read 15 119 Ω 49 (Post now available as a pdf: https://arxiv.org/abs/2212.11415 ) TL;DR: Unaligned AI will have a convergent instrumental incentive to make its thoughts difficult for us to interpret. In this article, I discuss many ways that a capable AI might circumvent scalable interpretability methods and

Explore this link on the map →

related reading