Faithful, Interpretable Model Explanations via Causal Abstraction | SAIL Blog
ai.stanford.edu · 4,028 words · saved by 3 readers
Seeking human-intelligible explanations
Seeking human-intelligible explanations Explaining why a deep learning model makes the predictions it does has emerged as one of the most challenging questions in AI ( Lipton 2018 , Pearl 2019 ). There is something of a paradox about this, however. After all, deep learning models are closed, deterministic systems that give us ground-truth knowledge of the causal relationships between all their components. Thus, their behavior is in many ways easy to explain: one can mechanistically walk through the mathematical operations or the associated computer code, and this can be done at varying levels
saved by
related reading
- Causal Scrubbing: a method for rigorously testing interpretability hypotheses [Redwood Research] — LessWronglesswrong.com
- A Comprehensive Mechanistic Interpretability Explainer & Glossary — Neel Nandaneelnanda.io
- The Hitchhiker's Guide to Actionable Interpretabilityactionable-interpretability-guide.github.io
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- The Building Blocks of Interpretabilitydistill.pub
- Transformer Circuits Threadtransformer-circuits.pub
- [2508.11214] How Causal Abstraction Underpins Computational Explanationarxiv.org
- Circuit Tracing: Revealing Computational Graphs in Language Modelstransformer-circuits.pub
- Topicslearnmechinterp.com
- How Causal Abstraction Underpins Computational Explanationarxiv.org
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- Language models can explain neurons in language modelsopenaipublic.blob.core.windows.net