✳flâneur — a map of the web's best reading
Neel Nanda on Mechanistic Interpretability: Progress, Limits, and Paths to Safer AI — EA Forum
forum.effectivealtruism.org · 9,250 words · saved by 3 readers
By Robert Wiblin | Watch on Youtube | Listen on Spotify | Transcript • ---------------------------------------- …
Neel Nanda on Mechanistic Interpretability: Progress, Limits, and Paths to Safer AI — EA Forum Hide table of contents The 80,000 Hours Podcast Neel Nanda on Mechanistic Interpretability: Progress, Limits, and Paths to Safer AI by 80000_Hours Sep 8 2025 37 min read 0 6 AI safety AI alignment AI interpretability DeepMind Podcasts Frontpage Neel Nanda on Mechanistic Interpretability: Progress, Limits, and Paths to Safer AI Episode summary The interview in a nutshell 1. Mech interp won’t solve alignment alone — but remains crucial 2. Simple techniques often outperform complex ones Probes beat fanc
Explore this link on the map →saved by
related reading
- Dario Amodei — The Urgency of Interpretabilitydarioamodei.com
- Neel Nanda on the race to read AI minds (part 1) | 80,000 Hours80000hours.org
- An Ambitious Vision for Interpretability — AI Alignment Forumalignmentforum.org
- A Pragmatic Vision for Interpretability — AI Alignment Forumalignmentforum.org
- How To Become A Mechanistic Interpretability Researcher — AI Alignment Forumalignmentforum.org
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- On Optimism for Interpretabilitygoodfire.ai
- How To Become A Mechanistic Interpretability Researcher — LessWronglesswrong.com
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- How Can Interpretability Researchers Help AGI Go Well? — AI Alignment Forumalignmentforum.org
- Why I'm Moving from Mechanistic to Prosaic Interpretability — LessWronglesswrong.com
- Against Almost Every Theory of Impact of Interpretability — AI Alignment Forumalignmentforum.org