✳flâneur — a map of the web's best reading
Faithful, Interpretable Model Explanations via Causal Abstraction | SAIL Blog
ai.stanford.edu · 4,028 words · saved by 3 readers
Seeking human-intelligible explanations
Seeking human-intelligible explanations Explaining why a deep learning model makes the predictions it does has emerged as one of the most challenging questions in AI ( Lipton 2018 , Pearl 2019 ). There is something of a paradox about this, however. After all, deep learning models are closed, deterministic systems that give us ground-truth knowledge of the causal relationships between all their components. Thus, their behavior is in many ways easy to explain: one can mechanistically walk through the mathematical operations or the associated computer code, and this can be done at varying levels
Explore this link on the map →saved by
related reading
- Causal Scrubbing: a method for rigorously testing interpretability hypotheses [Redwood Research] — LessWronglesswrong.com
- The Hitchhiker's Guide to Actionable Interpretabilityactionable-interpretability-guide.github.io
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- The Building Blocks of Interpretabilitydistill.pub
- Transformer Circuits Threadtransformer-circuits.pub
- CausaLab — Can LLM Agents Discover Causal Mechanisms by Experiment?dylanzsz.github.io
- Circuit Tracing: Revealing Computational Graphs in Language Modelstransformer-circuits.pub
- Dario Amodei — The Urgency of Interpretabilitydarioamodei.com
- Tracing the thoughts of a large language model \ Anthropicanthropic.com
- An Ambitious Vision for Interpretability — AI Alignment Forumalignmentforum.org
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- Language models can explain neurons in language modelsopenaipublic.blob.core.windows.net