Formal verification, heuristic explanations and surprise accounting — Alignment Research Center
ARC's current research focus can be thought of as trying to combine mechanistic interpretability and formal verification. If we had a deep understanding of what was going on inside a neural network, we would hope to be able to use that understanding to verify that the network was not going to behave dangerously in unforeseen situations. ARC is attempting to perform this kind of verification, but using a mathematical kind of "explanation" instead of one written in natural language. To help elucidate this connection, ARC has been supporting work on Compact Proofs of Model Performance via Mechanistic Interpretability by Jason Gross, Rajashree Agrawal, Lawrence Chan and others, which we were excited to see released along with this post. While we ultimately think that provable guarantees for large neural networks are unworkable as a long-term goal, we think that this work serves as a useful springboard towards alternatives. In this post, we will: We are also sharing a draft by Gabriel Wu (c
ARC's current research focus can be thought of as trying to combine mechanistic interpretability and formal verification. If we had a deep understanding of what was going on inside a neural network, we would hope to be able to use that understanding to verify that the network was not going to behave dangerously in unforeseen situations. ARC is attempting to perform this kind of verification, but using a mathematical kind of "explanation" instead of one written in natural language. To help elucidate this connection, ARC has been supporting work on Compact Proofs of Model Performance via Mechani
Explore this link on the map →saved by
related reading
- AlgZoo: uninterpreted models with fewer than 1,500 parameters — LessWronglesswrong.com
- ARC progress update: Competing with sampling — LessWronglesswrong.com
- Prediction, Explanation, or Over-interpretation?elena-baixy.github.io
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- Zoom In: An Introduction to Circuitsdistill.pub
- An Ambitious Vision for Interpretability — AI Alignment Forumalignmentforum.org
- Obstacles in ARC's agenda: Finding explanations — LessWronglesswrong.com
- Language models can explain neurons in language modelsopenaipublic.blob.core.windows.net
- Interpretability — LessWronglesswrong.com
- A Pragmatic Vision for Interpretability — AI Alignment Forumalignmentforum.org
- Competing with sampling — Alignment Research Centeralignment.org
- The Building Blocks of Interpretabilitydistill.pub