Mechanistic anomaly detection and ELK — LessWrong
(Follow-up to Eliciting Latent Knowledge. Describing joint work with Mark Xu. This is an informal description of ARC’s current research approach; not a polished product intended to be understandable to many people.) Suppose that I have a diamond in a vault, a collection of cameras, and an ML system that is excellent at predicting what those cameras will see over the next hour. I’d like to distinguish cases where the model predicts that the diamond will “actually” remain in the vault, from cases where the model predicts that someone will tamper with the cameras so that the diamond merely appears to remain in the vault. (Or cases where someone puts a fake diamond in its place, or…) One approach to this problem is to identify (the diamond remains in the vault) as the “normal” reason for the diamond to appear on camera. Then on a new input where the diamond appears on camera, we can ask whether it is for the normal reason or for a different reason. In this post I’ll describe an approach to
x Mechanistic anomaly detection and ELK — LessWrong Eliciting Latent Knowledge AI Curated 138 Mechanistic anomaly detection and ELK by paulfchristiano 25th Nov 2022 ai-alignment.com AI Alignment Forum 26 min read 22 138 Ω 67 ( Follow-up to Eliciting Latent Knowledge . Describing joint work with Mark Xu. This is an informal description of ARC’s current research approach; not a polished product intended to be understandable to many people. ) Suppose that I have a diamond in a vault, a collection of cameras, and an ML system that is excellent at predicting what those cameras will see over the nex
Explore this link on the map →related reading
- A gentle introduction to mechanistic anomaly detection — LessWronglesswrong.com
- A gentle introduction to mechanistic anomaly detection — LessWronglesswrong.com
- Concrete empirical research projects in mechanistic anomaly detection — LessWronglesswrong.com
- Mediumai-alignment.com
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- A bird's eye view of ARC's research — Alignment Research Centeralignment.org
- Neel Nanda on Mechanistic Interpretability: Progress, Limits, and Paths to Safer AI — EA Forumforum.effectivealtruism.org
- Eight Strategies for Tackling the Hard Part of the Alignment Problem — AI Alignment Forumalignmentforum.org
- A Mike's-Eye View of ARC's Research — Alignment Research Centeralignment.org
- 39 - Evan Hubinger on Model Organisms of Misalignment | AXRP - the AI X-risk Research Podcastaxrp.net
- Model Organisms of Misalignment: The Case for a New Pillar of Alignment Research — AI Alignment Forumalignmentforum.org
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org