flâneur — a map of the web's best reading

A gentle introduction to mechanistic anomaly detection — LessWrong

lesswrong.com · 3,898 words · saved by 2 readers

TL;DR: Mechanistic anomaly detection aims to flag when an AI produces outputs for “unusual reasons.” It is similar to mechanistic interpretability but doesn’t demand human understanding. I give a self-contained introduction to mechanistic anomaly detection from a slightly different angle than the existing one by Paul Christiano (focused less on heuristic arguments and drawing a more explicit parallel to interpretability). Mechanistic anomaly detection was first introduced by the Alignment Research Center (ARC), and a lot of this post is based on their ideas. However, I am not affiliated with ARC; this post represents my perspective. We want to create useful AI systems that never do anything too bad. Mechanistic anomaly detection relaxes this goal in two big ways: These are serious simplifications. But strong methods for mechanistic anomaly detection (or MAD for short) might still be important progress toward the full goal or even achieve it entirely: I intentionally say “unusual reason

x A gentle introduction to mechanistic anomaly detection — LessWrong Empirical mechanistic anomaly detection AI Frontpage 74 A gentle introduction to mechanistic anomaly detection by Erik Jenner 3rd Apr 2024 13 min read 2 74 TL;DR: Mechanistic anomaly detection aims to flag when an AI produces outputs for “unusual reasons.” It is similar to mechanistic interpretability but doesn’t demand human understanding. I give a self-contained introduction to mechanistic anomaly detection from a slightly different angle than the existing one by Paul Christiano (focused less on heuristic arguments and draw

Explore this link on the map →

saved by

related reading