✳flâneur — a map of the web's best reading
A gentle introduction to mechanistic anomaly detection — LessWrong
lesswrong.com · 3,901 words · saved by 1 readers
Mechanistic anomaly detection aims to flag when an AI produces outputs for “unusual reasons.” It is similar to mechanistic interpretability bu…
x A gentle introduction to mechanistic anomaly detection — LessWrong Empirical mechanistic anomaly detection AI Frontpage 74 A gentle introduction to mechanistic anomaly detection by Erik Jenner 3rd Apr 2024 13 min read 2 74 TL;DR: Mechanistic anomaly detection aims to flag when an AI produces outputs for “unusual reasons.” It is similar to mechanistic interpretability but doesn’t demand human understanding. I give a self-contained introduction to mechanistic anomaly detection from a slightly different angle than the existing one by Paul Christiano (focused less on heuristic arguments and draw
Explore this link on the map →related reading
- A gentle introduction to mechanistic anomaly detection — LessWronglesswrong.com
- Concrete empirical research projects in mechanistic anomaly detection — LessWronglesswrong.com
- Mechanistic anomaly detection and ELK — LessWronglesswrong.com
- Dario Amodei — The Urgency of Interpretabilitydarioamodei.com
- Neel Nanda on Mechanistic Interpretability: Progress, Limits, and Paths to Safer AI — EA Forumforum.effectivealtruism.org
- An Ambitious Vision for Interpretability — AI Alignment Forumalignmentforum.org
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- Eight Strategies for Tackling the Hard Part of the Alignment Problem — AI Alignment Forumalignmentforum.org
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- How To Become A Mechanistic Interpretability Researcher — AI Alignment Forumalignmentforum.org
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- A Pragmatic Vision for Interpretability — AI Alignment Forumalignmentforum.org