flâneur — a map of the web's best reading

Neel Nanda on Mechanistic Interpretability: Progress, Limits, and Paths to Safer AI — EA Forum

forum.effectivealtruism.org · 9,250 words · saved by 3 readers

By Robert Wiblin  |    Watch on Youtube   |   Listen on Spotify   |    Transcript • ---------------------------------------- …

Neel Nanda on Mechanistic Interpretability: Progress, Limits, and Paths to Safer AI — EA Forum Hide table of contents The 80,000 Hours Podcast Neel Nanda on Mechanistic Interpretability: Progress, Limits, and Paths to Safer AI by 80000_Hours Sep 8 2025 37 min read 0 6 AI safety AI alignment AI interpretability DeepMind Podcasts Frontpage Neel Nanda on Mechanistic Interpretability: Progress, Limits, and Paths to Safer AI Episode summary The interview in a nutshell 1. Mech interp won’t solve alignment alone — but remains crucial 2. Simple techniques often outperform complex ones Probes beat fanc

Explore this link on the map →

saved by

related reading