A gentle introduction to mechanistic anomaly detection — LessWrong
TL;DR: Mechanistic anomaly detection aims to flag when an AI produces outputs for “unusual reasons.” It is similar to mechanistic interpretability but doesn’t demand human understanding. I give a self-contained introduction to mechanistic anomaly detection from a slightly different angle than the existing one by Paul Christiano (focused less on heuristic arguments and drawing a more explicit parallel to interpretability). Mechanistic anomaly detection was first introduced by the Alignment Research Center (ARC), and a lot of this post is based on their ideas. However, I am not affiliated with ARC; this post represents my perspective. We want to create useful AI systems that never do anything too bad. Mechanistic anomaly detection relaxes this goal in two big ways: These are serious simplifications. But strong methods for mechanistic anomaly detection (or MAD for short) might still be important progress toward the full goal or even achieve it entirely: I intentionally say “unusual reason
x A gentle introduction to mechanistic anomaly detection — LessWrong Empirical mechanistic anomaly detection AI Frontpage 74 A gentle introduction to mechanistic anomaly detection by Erik Jenner 3rd Apr 2024 13 min read 2 74 TL;DR: Mechanistic anomaly detection aims to flag when an AI produces outputs for “unusual reasons.” It is similar to mechanistic interpretability but doesn’t demand human understanding. I give a self-contained introduction to mechanistic anomaly detection from a slightly different angle than the existing one by Paul Christiano (focused less on heuristic arguments and draw
Explore this link on the map →saved by
related reading
- A gentle introduction to mechanistic anomaly detection — LessWronglesswrong.com
- Concrete empirical research projects in mechanistic anomaly detection — LessWronglesswrong.com
- Mechanistic anomaly detection and ELK — LessWronglesswrong.com
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- Dario Amodei — The Urgency of Interpretabilitydarioamodei.com
- Neel Nanda on Mechanistic Interpretability: Progress, Limits, and Paths to Safer AI — EA Forumforum.effectivealtruism.org
- An Ambitious Vision for Interpretability — AI Alignment Forumalignmentforum.org
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- Eight Strategies for Tackling the Hard Part of the Alignment Problem — AI Alignment Forumalignmentforum.org
- How To Become A Mechanistic Interpretability Researcher — AI Alignment Forumalignmentforum.org
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- A Pragmatic Vision for Interpretability — AI Alignment Forumalignmentforum.org