How To Become A Mechanistic Interpretability Researcher — AI Alignment Forum
Note: If you’ll forgive the shameless self-promotion, applications for my MATS stream are open until Sept 12. I help people write a mech interp paper, often accept promising people new to mech interp, and alumni often have careers as mech interp researchers. If you’re interested in this post I recommend applying! The application should be educational whatever happens: you spend a weekend doing a small mech interp research project, and show me what you learned. Last updated Sept 2 2025 Mechanistic interpretability (mech interp) is, in my incredibly biased opinion, one of the most exciting research areas out there. We have these incredibly complex AI models that we don't understand, yet there are tantalizing signs of real structure inside them. Even partial understanding of this structure opens up a world of possibilities, yet is neglected by 99% of machine learning researchers. There’s so much to do! I think mech interp is an unusually easy field to learn about on your own: there’s a lo
x How To Become A Mechanistic Interpretability Researcher — AI Alignment Forum Careers Interpretability (ML & AI) Scholarship & Learning Practical AI Frontpage 2025 Top Fifty: 50 % 36 How To Become A Mechanistic Interpretability Researcher by Neel Nanda 2nd Sep 2025 67 min read 12 36 Last updated Sept 2 2025 Note - if you want to pursue a career in this kind of research, apply to my MATS stream! Apps aren't currently open, sign up here to be notified TL;DR This post is about the mindset and process I recommend if you want to do mechanistic interpretability research. I aim to give a clear sense
Explore this link on the map →saved by
related reading
- How To Become A Mechanistic Interpretability Researcher — LessWronglesswrong.com
- A Pragmatic Vision for Interpretability — AI Alignment Forumalignmentforum.org
- Neel Nanda on Mechanistic Interpretability: Progress, Limits, and Paths to Safer AI — EA Forumforum.effectivealtruism.org
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- An Ambitious Vision for Interpretability — AI Alignment Forumalignmentforum.org
- Dario Amodei — The Urgency of Interpretabilitydarioamodei.com
- Why I'm Moving from Mechanistic to Prosaic Interpretability — LessWronglesswrong.com
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- EIS XIV: Is mechanistic interpretability about to be practically useful? — AI Alignment Forumalignmentforum.org
- A Pragmatic Vision for Interpretability — LessWronglesswrong.com
- How Can Interpretability Researchers Help AGI Go Well? — AI Alignment Forumalignmentforum.org
- The Hitchhiker's Guide to Actionable Interpretabilityactionable-interpretability-guide.github.io