Mech interp is not pre-paradigmatic — LessWrong
So, as a field, we don't have to be happy with the dominant paradigm. But just because we're not happy with it doesn't mean it's not there. Um, ok fine, so what alternative term do you propose to replace "pre-paradigmatic" as it is currently used, to indicate that there's no remotely satisfactory paradigm in which to get going on the parts of the field-to-be that really matter? This is a blogpost version of a talk I gave earlier this year at GDM. Epistemic status: Vague and handwavy. Nuance is often missing. Some of the claims depend on implicit definitions that may be reasonable to disagree with and is, in an important sense, subjective. But overall I think it's directionally (subjectively) true. It's often said that mech interp is pre-paradigmatic. I think it's worth being skeptical of this claim. In this post I argue that: First, we need to be familiar with the basic definition of a paradigm: A paradigm is a distinct set of concepts or thought patterns, including theories, resea
x Mech interp is not pre-paradigmatic — LessWrong Interpretability (ML & AI) AI Frontpage 2025 Top Fifty: 9 % 213 Mech interp is not pre-paradigmatic by Lee Sharkey 10th Jun 2025 AI Alignment Forum 16 min read 15 213 Ω 81 This is a blogpost version of a talk I gave earlier this year at GDM. Epistemic status: Vague and handwavy. Nuance is often missing. Some of the claims depend on implicit definitions that may be reasonable to disagree with and is, in an important sense, subjective. But overall I think it's directionally (subjectively) true. It's often said that mech interp is pre-paradigmatic
Explore this link on the map →related reading
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- An Ambitious Vision for Interpretability — AI Alignment Forumalignmentforum.org
- How To Become A Mechanistic Interpretability Researcher — AI Alignment Forumalignmentforum.org
- Neel Nanda on Mechanistic Interpretability: Progress, Limits, and Paths to Safer AI — EA Forumforum.effectivealtruism.org
- A Pragmatic Vision for Interpretability — AI Alignment Forumalignmentforum.org
- How To Become A Mechanistic Interpretability Researcher — LessWronglesswrong.com
- Interpretability Dreamstransformer-circuits.pub
- Dario Amodei — The Urgency of Interpretabilitydarioamodei.com
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- Sparsify: A mechanistic interpretability research agenda — AI Alignment Forumalignmentforum.org
- Interpreting Language Model Parametersgoodfire.ai
- Why I'm Moving from Mechanistic to Prosaic Interpretability — LessWronglesswrong.com