flâneur — a map of the web's best reading

Interpretability Will Not Reliably Find Deceptive AI — AI Alignment Forum

alignmentforum.org · 3,829 words · saved by 1 readers

Disclaimer: Post written in a personal capacity. These are personal hot takes and do not in any way represent my employer's views • TL;DR: …

x Interpretability Will Not Reliably Find Deceptive AI — AI Alignment Forum Interpretability (ML & AI) AI Curated 2025 Top Fifty: 38 % 117 Interpretability Will Not Reliably Find Deceptive AI by Neel Nanda 4th May 2025 9 min read 69 117 Disclaimer: Post written in a personal capacity. These are personal opinions and do not in any way represent my employer's views TL;DR: I do not think we will produce high reliability methods to evaluate or monitor the safety of superintelligent systems via current research paradigms, with interpretability or otherwise. Interpretability still seems a valuable t

Explore this link on the map →

related reading