EIS XIV: Is mechanistic interpretability about to be practically useful? — AI Alignment Forum
Lately, I have been thinking of interpretability research as falling into five different tiers of rigor. This is when researchers claim they have succeeded in interpreting a model by definition or based on analyzing results and asserting hypotheses about them. This is a key part of the scientific method. But by itself, it is not good science. Previously in this sequence, I have argued that this standard is fairly pervasive. This is when researchers develop an interpretation, use it to make some (usually simple) prediction, and then show that this prediction validates. This is at least doing science, but it doesn't necessarily demonstrate any usefulness or value. This is when researchers accomplish a useful type of task with an interpretability technique but do so in a way that is toy, cherry-picked, or under a streetlight. This is when researchers show that an interpretability tool can be used uniquely or competitively to accomplish a useful task. For this level of rigor, it needs to b
x EIS XIV: Is mechanistic interpretability about to be practically useful? — AI Alignment Forum The Engineer’s Interpretability Sequence Interpretability (ML & AI) AI Frontpage 31 EIS XIV: Is mechanistic interpretability about to be practically useful? by scasper 11th Oct 2024 8 min read 4 31 Part 14 of 12 in the Engineer’s Interpretability Sequence . Is this market really only at 63%? I think you should take the over. Only 63%? I think you should take the over. Five tiers of rigor for safety-oriented interpretability work Lately, I have been thinking of interpretability research as falling in
Explore this link on the map →related reading
- Dario Amodei — The Urgency of Interpretabilitydarioamodei.com
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- How To Become A Mechanistic Interpretability Researcher — AI Alignment Forumalignmentforum.org
- Neel Nanda on Mechanistic Interpretability: Progress, Limits, and Paths to Safer AI — EA Forumforum.effectivealtruism.org
- An Ambitious Vision for Interpretability — AI Alignment Forumalignmentforum.org
- How To Become A Mechanistic Interpretability Researcher — LessWronglesswrong.com
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- A Pragmatic Vision for Interpretability — AI Alignment Forumalignmentforum.org
- The Hitchhiker's Guide to Actionable Interpretabilityactionable-interpretability-guide.github.io
- Why I'm Moving from Mechanistic to Prosaic Interpretability — LessWronglesswrong.com
- On Optimism for Interpretabilitygoodfire.ai
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumneelnanda.io