flâneur — a map of the web's best reading

Tracing Attention Computation Through Feature Interactions

transformer-circuits.pub · 2,932 words · saved by 1 readers

Transformer-based language models involve two main kinds of computations: multi-layer perceptron (MLP) layers that process information within a context position, and attention layers that conditionally move and process information between context positions. In our recent papers we made significant progress in breaking down MLP computation into interpretable steps. In this update, we fill in a major missing piece in our methodology, by introducing a way to decompose attentional computations as well.

Tracing Attention Computation Through Feature Interactions Transformer Circuits Thread Tracing Attention Computation Through Feature Interactions We describe and apply a method to explain attention patterns in terms of feature interactions, and integrate this information into attribution graphs. Authors Harish Kamath * , Emmanuel Ameisen * , Isaac Kauvar, Rodrigo Luger, Wes Gurnee, Adam Pearce, Sam Zimmerman, Joshua Batson, Thomas Conerly, Chris Olah, Jack Lindsey ‡ Affiliations Anthropic Published July 31st, 2025 * Core Research Contributor; ‡ Correspondence to jacklindsey@anthropic.com Trans

Explore this link on the map →

related reading