Sparsify: A mechanistic interpretability research agenda — AI Alignment Forum
Over the last couple of years, mechanistic interpretability has seen substantial progress. Part of this progress has been enabled by the identification of superposition as a key barrier to understanding neural networks (Elhage et al., 2022) and the identification of sparse autoencoders as a solution to superposition (Sharkey et al., 2022; Cunningham et al., 2023; Bricken et al., 2023). From our current vantage point, I think there’s a relatively clear roadmap toward a world where mechanistic interpretability is useful for safety. This post outlines my views on what progress in mechanistic interpretability looks like and what I think is achievable by the field in the next 2+ years. It represents a rough outline of what I plan to work on in the near future. My thinking and work is, of course, very heavily inspired by the work of Chris Olah, other Anthropic researchers, and other early mechanistic interpretability researchers. In addition to sharing some personal takes, this article bring
x Sparsify: A mechanistic interpretability research agenda — AI Alignment Forum Interpretability (ML & AI) Sparse Autoencoders (SAEs) Research Agendas Apollo Research (org) Practice & Philosophy of Science AI Frontpage 43 Sparsify: A mechanistic interpretability research agenda by Lee Sharkey 3rd Apr 2024 27 min read 23 43 Over the last couple of years, mechanistic interpretability has seen substantial progress. Part of this progress has been enabled by the identification of superposition as a key barrier to understanding neural networks ( Elhage et al., 2022 ) and the identification of sparse
Explore this link on the map →related reading
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- An Ambitious Vision for Interpretability — AI Alignment Forumalignmentforum.org
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- Towards Monosemanticity: Decomposing Language Models With Dictionary Learningtransformer-circuits.pub
- An Intuitive Explanation of Sparse Autoencoders for LLM Interpretability | Adam Karvonenadamkarvonen.github.io
- Introduction to Mechanistic Interpretability - by Sarahblog.bluedot.org
- Dario Amodei — The Urgency of Interpretabilitydarioamodei.com
- A List of 45+ Mech Interp Project Ideas from Apollo Research’s Interpretability Team — AI Alignment Forumalignmentforum.org
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumneelnanda.io
- Neel Nanda on Mechanistic Interpretability: Progress, Limits, and Paths to Safer AI — EA Forumforum.effectivealtruism.org
- A gentle introduction to sparse autoencoders — LessWronglesswrong.com
- A Pragmatic Vision for Interpretability — AI Alignment Forumalignmentforum.org