[2404.15255] How to use and interpret activation patching
Abstract:Activation patching is a popular mechanistic interpretability technique, but has many subtleties regarding how it is applied and how one may interpret the results. We provide a summary of advice and best practices, based on our experience using this technique in practice. We include an overview of the different ways to apply activation patching and a discussion on how to interpret the results. We focus on what evidence patching experiments provide about circuits, and on the choice of metric and associated pitfalls.
How to use and interpret activation patching Stefan Heimersheim Neel Nanda stefan.heimersheim@gmail.com arXiv:2404.15255v1 [cs.LG] 23 Apr 2024 Abstract Activation patching is a popular mechanistic interpretability technique, but has many subtleties…
related reading
- Attribution Patching: Activation Patching At Industrial Scale - Neel Nandaneelnanda.io
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- Towards Automated Circuit Discovery for Mechanistic Interpretabilityarxiv.org
- Circuit Tracing: Revealing Computational Graphs in Language Modelstransformer-circuits.pub
- Causal Scrubbing: a method for rigorously testing interpretability hypotheses [Redwood Research] — LessWronglesswrong.com
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumneelnanda.io
- The Building Blocks of Interpretabilitydistill.pub
- Faithful, Interpretable Model Explanations via Causal Abstraction | SAIL Blogai.stanford.edu
- Topicslearnmechinterp.com
- An Ambitious Vision for Interpretability — AI Alignment Forumalignmentforum.org
- Attribution-based parameter decomposition — LessWronglesswrong.com