flâneur — a map of the web's best reading

Concrete empirical research projects in mechanistic anomaly detection — LessWrong

lesswrong.com · 3,386 words · saved by 1 readers

Thanks to Jordan Taylor, Mark Xu, Alex Mallen, and Lawrence Chan for feedback on a draft! This post was mostly written by Erik, but we're all currently collaborating on this research direction. Mechanistic anomaly detection (MAD) aims to flag when an AI produces outputs for “unusual reasons.” It is similar to mechanistic interpretability but doesn’t demand human understanding. The Alignment Research Center (ARC) is trying to formalize “reasons” for an AI’s output using heuristic arguments, aiming for an indefinitely scalable solution to MAD. As a complement to ARC’s theoretical approach, we are excited about empirical research on MAD. Rather than looking for a principled definition of “reasons,” this means creating incrementally harder MAD benchmarks and better MAD methods. We have been thinking about and working on empirical MAD research for the past months. We believe there are many tractable and useful experiments, only a fraction of which we can run ourselves. This post describes s

x Concrete empirical research projects in mechanistic anomaly detection — LessWrong Empirical mechanistic anomaly detection Open Problems AI Frontpage 43 Concrete empirical research projects in mechanistic anomaly detection by Erik Jenner , Viktor Rehnberg , Oliver Daniels 3rd Apr 2024 12 min read 3 43 Thanks to Jordan Taylor, Mark Xu, Alex Mallen, and Lawrence Chan for feedback on a draft! This post was mostly written by Erik, but we're all currently collaborating on this research direction. Mechanistic anomaly detection (MAD) aims to flag when an AI produces outputs for “unusual reasons.” It

Explore this link on the map →

related reading