What is the purpose of interpretability?
I gave a five-minute lightning talk at the Mechanistic Interpretability Workshop at ICML this year in Seoul. Many thanks to the organizers, who encouraged us to speak about our "high-level vision" and "what the field should be doing differently". My talk ended up being indeed quite high-level and personal. I've reproduced a slightly expanded version of it below. Today I'll share some very high-level thoughts on my hopes for the field of interpretability—what we could become, or fail to become, in the coming years. Many of these thoughts are what come up for me when reflecting on the call for "pragmatic interpretability" which was of course the topic of so much discussion at this workshop at NeurIPS in December.1 I'll start with something I believe, which is that the task of understanding neural networks is a very deep science. I believe that if we froze AI progress, we could study today's models for decades, possibly centuries, and that that project would become as culturally significa
What is the purpose of interpretability? What is the purpose of interpretability? 2026-07-13 I gave a five-minute lightning talk at the Mechanistic Interpretability Workshop at ICML this year in Seoul. Many thanks to the organizers, who encouraged us to speak about our "high-level vision" and "what the field should be doing differently". My talk ended up being indeed quite high-level and personal. I've reproduced a slightly expanded version of it below. Today I'll share some very high-level thoughts on my hopes for the field of interpretability—what we could become, or fail to become, in the c
Explore this link on the map →related reading
- Dario Amodei — The Urgency of Interpretabilitydarioamodei.com
- On Optimism for Interpretabilitygoodfire.ai
- An Ambitious Vision for Interpretability — AI Alignment Forumalignmentforum.org
- A Pragmatic Vision for Interpretability — AI Alignment Forumalignmentforum.org
- A short note on interpretability and mindsericjmichaud.com
- Intentionally Designing the Future of AIgoodfire.ai
- Neel Nanda on Mechanistic Interpretability: Progress, Limits, and Paths to Safer AI — EA Forumforum.effectivealtruism.org
- Transformer Circuits Threadtransformer-circuits.pub
- The Building Blocks of Interpretabilitydistill.pub
- The Hitchhiker's Guide to Actionable Interpretabilityactionable-interpretability-guide.github.io
- On Developing a Mathematical Theory of Interpretability — LessWronglesswrong.com
- Interpretability — LessWronglesswrong.com