flâneur — a map of the web's best reading

How To Become A Mechanistic Interpretability Researcher — AI Alignment Forum

alignmentforum.org · 17,549 words · saved by 5 readers

Note: If you’ll forgive the shameless self-promotion, applications for my MATS stream are open until Sept 12. I help people write a mech interp paper, often accept promising people new to mech interp, and alumni often have careers as mech interp researchers. If you’re interested in this post I recommend applying! The application should be educational whatever happens: you spend a weekend doing a small mech interp research project, and show me what you learned. Last updated Sept 2 2025 Mechanistic interpretability (mech interp) is, in my incredibly biased opinion, one of the most exciting research areas out there. We have these incredibly complex AI models that we don't understand, yet there are tantalizing signs of real structure inside them. Even partial understanding of this structure opens up a world of possibilities, yet is neglected by 99% of machine learning researchers. There’s so much to do! I think mech interp is an unusually easy field to learn about on your own: there’s a lo

x How To Become A Mechanistic Interpretability Researcher — AI Alignment Forum Careers Interpretability (ML & AI) Scholarship & Learning Practical AI Frontpage 2025 Top Fifty: 50 % 36 How To Become A Mechanistic Interpretability Researcher by Neel Nanda 2nd Sep 2025 67 min read 12 36 Last updated Sept 2 2025 Note - if you want to pursue a career in this kind of research, apply to my MATS stream! Apps aren't currently open, sign up here to be notified TL;DR This post is about the mindset and process I recommend if you want to do mechanistic interpretability research. I aim to give a clear sense

Explore this link on the map →

saved by

related reading