Attribution Patching: Activation Patching At Industrial Scale — Neel Nanda
A write-up of an incomplete project I worked on at Anthropic in early 2022, using gradient-based approximation to make activation patching far more scalable
Attribution Patching: Activation Patching At Industrial Scale Feb 4 Written By Neel Nanda The following is a write-up of an (incomplete) project I worked on while at Anthropic, and a significant amount of the credit goes to the then team, Chris Olah, Catherine Olsson, Nelson Elhage & Tristan Hume. I've since cleaned up this project in my personal time and personal capacity. TLDR Activation patching is an existing technique for identifying which model activations are most important for determining model behaviour between two similar prompts that differ in a key detail I introduce a technique ca
Explore this link on the map →saved by
related reading
- Circuit Tracing: Revealing Computational Graphs in Language Modelstransformer-circuits.pub
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- Causal Scrubbing: a method for rigorously testing interpretability hypotheses [Redwood Research] — LessWronglesswrong.com
- Transformer Circuits Threadtransformer-circuits.pub
- A Mathematical Framework for Transformer Circuitstransformer-circuits.pub
- Chapter 1: Transformer Interpretability - ARENAlearn.arena.education
- On the Biology of a Large Language Modeltransformer-circuits.pub
- Weight-Sparse Circuits May Be Interpretable Yet Unfaithful — LessWronglesswrong.com
- Interpretability Infrastructure at Frontier Scale: Harvesting Activations from a Trillion-Parameter Modelgoodfire.ai
- Chapter 1: Transformer Interpretability - ARENAlearn.arena.education
- [2404.15255] How to use and interpret activation patchingarxiv.org