The Building Blocks of Interpretability
Interpretability techniques are normally studied in isolation. We explore the powerful interfaces that arise when you combine them -- and the rich structure of this combinatorial space.
The Building Blocks of Interpretability Distill The Building Blocks of Interpretability Developing a Grammar of Interfaces for Understanding Neural Networks --> Interpretability techniques are normally studied in isolation. We explore the powerful interfaces that arise when you combine them — and the rich structure of this combinatorial space. Authors Affiliations Chris Olah Google Brain Arvind Satyanarayan Google Brain Ian Johnson Google Cloud Shan Carter Google Brain Ludwig Schubert Google Brain Katherine Ye CMU Alexander Mordvintsev Google Research Published March 6, 2018 DOI 10.23915/disti
saved by
related reading
- Toy Models of Superpositiontransformer-circuits.pub
- A Comprehensive Mechanistic Interpretability Explainer & Glossary — Neel Nandaneelnanda.io
- Feature Visualizationdistill.pub
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- Towards Monosemanticity: Decomposing Language Models With Dictionary Learningtransformer-circuits.pub
- Softmax Linear Unitstransformer-circuits.pub
- Towards Monosemanticity: Decomposing Language Models With Dictionary Learningtransformer-circuits.pub
- Faithful, Interpretable Model Explanations via Causal Abstraction | SAIL Blogai.stanford.edu
- Towards Automated Circuit Discovery for Mechanistic Interpretabilityarxiv.org
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- Transformer Circuits Threadtransformer-circuits.pub
- Activation space interpretability may be doomed — LessWronglesswrong.com