Mech interp is not pre-paradigmatic — LessWrong
So, as a field, we don't have to be happy with the dominant paradigm. But just because we're not happy with it doesn't mean it's not there. Um, ok fine, so what alternative term do you propose to replace "pre-paradigmatic" as it is currently used, to indicate that there's no remotely satisfactory paradigm in which to get going on the parts of the field-to-be that really matter? This is a blogpost version of a talk I gave earlier this year at GDM. Epistemic status: Vague and handwavy. Nuance is often missing. Some of the claims depend on implicit definitions that may be reasonable to disagree with and is, in an important sense, subjective. But overall I think it's directionally (subjectively) true. It's often said that mech interp is pre-paradigmatic. I think it's worth being skeptical of this claim. In this post I argue that: First, we need to be familiar with the basic definition of a paradigm: A paradigm is a distinct set of concepts or thought patterns, including theories, resea
x Mech interp is not pre-paradigmatic — LessWrong Interpretability (ML & AI) AI Frontpage 2025 Top Fifty: 9 % 213 Mech interp is not pre-paradigmatic by Lee Sharkey 10th Jun 2025 AI Alignment Forum 16 min read 15 213 Ω 81 This is a blogpost version of a talk I gave earlier this year at GDM. Epistemic status: Vague and handwavy. Nuance is often missing. Some of the claims depend on implicit definitions that may be reasonable to disagree with and is, in an important sense, subjective. But overall I think it's directionally (subjectively) true. It's often said that mech interp is pre-paradigmatic
related reading
- How To Become A Mechanistic Interpretability Researcher — AI Alignment Forumalignmentforum.org
- An Ambitious Vision for Interpretability — AI Alignment Forumalignmentforum.org
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- A Pragmatic Vision for Interpretability — AI Alignment Forumalignmentforum.org
- Neel Nanda on Mechanistic Interpretability: Progress, Limits, and Paths to Safer AI — EA Forumforum.effectivealtruism.org
- Interpretability in Parameter Space: Minimizing Mechanistic Description Length with Attribution-based Parameter Decompositionarxiv.org
- Interpretability Dreamstransformer-circuits.pub
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- Mechanistic?arxiv.org
- How To Become A Mechanistic Interpretability Researcher — LessWronglesswrong.com
- A Comprehensive Mechanistic Interpretability Explainer & Glossary — Neel Nandaneelnanda.io
- A short note on interpretability and mindsericjmichaud.com