✳flâneur — a map of the web's best reading
How To Become A Mechanistic Interpretability Researcher — LessWrong
lesswrong.com · 19,411 words · saved by 2 readers
Last updated Sept 2 2025 • Note - if you want to pursue a career in this kind of research, apply to my MATS stream! Due Dec 23 …
x How To Become A Mechanistic Interpretability Researcher — LessWrong Careers Interpretability (ML & AI) Scholarship & Learning Practical AI Frontpage 2025 Top Fifty: 50 % 146 How To Become A Mechanistic Interpretability Researcher by Neel Nanda 2nd Sep 2025 AI Alignment Forum 67 min read 12 146 Ω 36 Last updated Sept 2 2025 Note - if you want to pursue a career in this kind of research, apply to my MATS stream! Apps aren't currently open, sign up here to be notified TL;DR This post is about the mindset and process I recommend if you want to do mechanistic interpretability research. I aim to g
Explore this link on the map →saved by
related reading
- How To Become A Mechanistic Interpretability Researcher — AI Alignment Forumalignmentforum.org
- A Pragmatic Vision for Interpretability — AI Alignment Forumalignmentforum.org
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- Neel Nanda on Mechanistic Interpretability: Progress, Limits, and Paths to Safer AI — EA Forumforum.effectivealtruism.org
- An Ambitious Vision for Interpretability — AI Alignment Forumalignmentforum.org
- Transformer Circuits Threadtransformer-circuits.pub
- Dario Amodei — The Urgency of Interpretabilitydarioamodei.com
- Why I'm Moving from Mechanistic to Prosaic Interpretability — LessWronglesswrong.com
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- A Pragmatic Vision for Interpretability — LessWronglesswrong.com
- EIS XIV: Is mechanistic interpretability about to be practically useful? — AI Alignment Forumalignmentforum.org
- The Hitchhiker's Guide to Actionable Interpretabilityactionable-interpretability-guide.github.io