[2410.09087] Mechanistic?
Abstract:The rise of the term "mechanistic interpretability" has accompanied increasing interest in understanding neural models -- particularly language models. However, this jargon has also led to a fair amount of confusion. So, what does it mean to be "mechanistic"? We describe four uses of the term in interpretability research. The most narrow technical definition requires a claim of causality, while a broader technical definition allows for any exploration of a model's internals. However, the term also has a narrow cultural definition describing a cultural movement. To understand this semantic drift, we present a history of the NLP interpretability community and the formation of the separate, parallel "mechanistic" interpretability community. Finally, we discuss the broad cultural definition -- encompassing the entire field of interpretability -- and why the traditional NLP interpretability community has come to embrace it. We argue that the polysemy of "mechanistic" is the product of a critical divide within the interpretability community.
Title:Mechanistic? View PDF HTML (experimental) Abstract:The rise of the term "mechanistic interpretability" has accompanied increasing interest in understanding neural models -- particularly language models. However, this jargon has also led to a fair amount of confusion. So, what does it mean to be "mechanistic"? We describe four uses of the term in interpretability research. The most narrow technical definition requires a claim of causality, while a broader technical definition allows for any exploration of a model's internals. However, the term also has a narrow cultural definition…
saved by
related reading
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- How To Become A Mechanistic Interpretability Researcher — AI Alignment Forumalignmentforum.org
- Dario Amodei — The Urgency of Interpretabilitydarioamodei.com
- Interpretability in Parameter Space: Minimizing Mechanistic Description Length with Attribution-based Parameter Decompositionarxiv.org
- Towards Automated Circuit Discovery for Mechanistic Interpretabilityarxiv.org
- A Comprehensive Mechanistic Interpretability Explainer & Glossary — Neel Nandaneelnanda.io
- A Pragmatic Vision for Interpretability — LessWronglesswrong.com
- Assessing skeptical views of interpretability research | Christopher Pottsweb.stanford.edu
- Causal Scrubbing: a method for rigorously testing interpretability hypotheses [Redwood Research] — LessWronglesswrong.com
- An Ambitious Vision for Interpretability — AI Alignment Forumalignmentforum.org
- Neel Nanda on Mechanistic Interpretability: Progress, Limits, and Paths to Safer AI — EA Forumforum.effectivealtruism.org