A Toy Model of Interference Weights
An informal note on interference weights by Chris Olah, Nicholas L Turner, and Tom Conerly. Published July 29th, 2025. Not published yet. No DOI yet. This note explores the phenomenon of "interference weights" and "weight superposition", an idea that we've discussed briefly in previous papers and updates. We've come to believe they are a central issue if one wants to move from attribution graphs which describe why the model behaves the way it does for a specific example, to global circuit analysis where one can reason about the model more broadly. (In fact, avoiding interference weights was our primary motivation for studying attribution graphs.) We study interference weights in the context of toy models and preliminarily find: Key Takeaway 1: Interference weights can be demonstrated in toy models. We just need to slightly modify our interpretation of the setup from the original Toy Models paper. The resulting interference weights exhibit distinctive phenomenology we saw in Towards Mon
A Toy Model of Interference Weights Transformer Circuits Thread A Toy Model of Interference Weights An informal note on interference weights by Chris Olah, Nicholas L Turner, and Tom Conerly. Published July 29th, 2025. This note explores the phenomenon of "interference weights" and "weight superposition", an idea that we've discussed briefly in previous papers and updates. We've come to believe they are a central issue if one wants to move from attribution graphs which describe why the model behaves the way it does for a specific example, to global circuit analysis where one can reason about t
Explore this link on the map →related reading
- Toy Models of Superpositiontransformer-circuits.pub
- Interpretability Dreamstransformer-circuits.pub
- Zoom In: An Introduction to Circuitsdistill.pub
- Circuit Tracing: Revealing Computational Graphs in Language Modelstransformer-circuits.pub
- The Waluigi Effect (mega-post) — LessWronglesswrong.com
- They're Made Out of Weightsmaxleiter.com
- Inkling: Our Open-Weights Model - Thinking Machines Labthinkingmachines.ai
- Circuits Updates — May 2023transformer-circuits.pub
- What Would Non-Linear Features Actually Look Like? — Liv Gortonlivgorton.com
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- Weight-Sparse Circuits May Be Interpretable Yet Unfaithful — LessWronglesswrong.com
- Computational Superposition in a Toy Model of the U-AND Problem — LessWronglesswrong.com