Giang Nguyen on X: "Time for a mech interpretability conference?" / X
x.com · 1,284 words · saved by 1 readers
https://t.co/F0GteaVjQs
Why I wrote this? A workshop was harder to get into than the conference that hosts it (~24% for workshop < ~26% for ICML main conf). Several days ago, @BlancheMinerva posted that "ICML Mech Interp had a lower acceptance rate than ICML this year." Mechanistic interpretability, or mech interp for short, is the study of reverse-engineering the internal computations of neural networks. An ICML workshop is more selective than ICML main itself, how??!! I did a double check. FWIW, a workshop is meant to be the exploratory, low-stakes corner of a conference. How does it end up tighter than the…
related reading
- How To Become A Mechanistic Interpretability Researcher — AI Alignment Forumalignmentforum.org
- Neel Nanda on Mechanistic Interpretability: Progress, Limits, and Paths to Safer AI — EA Forumforum.effectivealtruism.org
- A Pragmatic Vision for Interpretability — AI Alignment Forumalignmentforum.org
- How To Become A Mechanistic Interpretability Researcher — LessWronglesswrong.com
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- Dario Amodei — The Urgency of Interpretabilitydarioamodei.com
- Assessing skeptical views of interpretability research | Christopher Pottsweb.stanford.edu
- An Ambitious Vision for Interpretability — AI Alignment Forumalignmentforum.org
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- A Pragmatic Vision for Interpretability — LessWronglesswrong.com
- Why I'm Moving from Mechanistic to Prosaic Interpretability — LessWronglesswrong.com
- Mechanistic?arxiv.org