The Building Blocks of Interpretability
Interpretability techniques are normally studied in isolation. We explore the powerful interfaces that arise when you combine them -- and the rich structure of this combinatorial space.
The Building Blocks of Interpretability Distill The Building Blocks of Interpretability Developing a Grammar of Interfaces for Understanding Neural Networks --> Interpretability techniques are normally studied in isolation. We explore the powerful interfaces that arise when you combine them — and the rich structure of this combinatorial space. Authors Affiliations Chris Olah Google Brain Arvind Satyanarayan Google Brain Ian Johnson Google Cloud Shan Carter Google Brain Ludwig Schubert Google Brain Katherine Ye CMU Alexander Mordvintsev Google Research Published March 6, 2018 DOI 10.23915/disti
Explore this link on the map →saved by
related reading
- Toy Models of Superpositiontransformer-circuits.pub
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- Towards Monosemanticity: Decomposing Language Models With Dictionary Learningtransformer-circuits.pub
- Softmax Linear Unitstransformer-circuits.pub
- Feature Visualizationdistill.pub
- Faithful, Interpretable Model Explanations via Causal Abstraction | SAIL Blogai.stanford.edu
- The Hitchhiker's Guide to Actionable Interpretabilityactionable-interpretability-guide.github.io
- Activation space interpretability may be doomed — LessWronglesswrong.com
- Transformer Circuits Threadtransformer-circuits.pub
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- Dario Amodei — The Urgency of Interpretabilitydarioamodei.com
- Distill — Latest articles about machine learningdistill.pub