flâneur — a map of the web's best reading

Scalable End-to-End Interpretability — LessWrong

lesswrong.com · 1,004 words · saved by 1 readers

This is partly a linkpost for Predictive Concept Decoders, and partly a response to Neel Nanda's Pragmatic Vision for AI Interpretability and Leo Gao…

x Scalable End-to-End Interpretability — LessWrong Interpretability (ML & AI) AI Frontpage 2025 Top Fifty: 14 % 121 Scalable End-to-End Interpretability by jsteinhardt 18th Dec 2025 AI Alignment Forum 3 min read 3 121 Ω 56 This is partly a linkpost for Predictive Concept Decoders , and partly a response to Neel Nanda's Pragmatic Vision for AI Interpretability and Leo Gao's Ambitious Vision for Interpretability . There is currently somewhat of a debate in the interpretability community between pragmatic interpretability ---grounding problems in empirically measurable safety tasks---and ambitiou

Explore this link on the map →

related reading