Stage-Wise Model Diffing
We report some developing work on the Anthropic interpretability team, which might be of interest to researchers working actively in this space. We'd ask you to treat these results like those of a colleague sharing some thoughts or preliminary experiments for a few minutes at a lab meeting, rather than a mature paper. This work presents a novel approach to "model diffing" with dictionary learning that reveals how the features of a transformer change from finetuning. This approach takes an initial sparse autoencoder (SAE) dictionary trained on the transformer before it has been finetuned, and finetunes the dictionary itself on either the new finetuning dataset or the finetuned transformer model. By tracking how dictionary features evolve through the different fine-tunes, we can isolate the effects of both dataset and model changes. We demonstrate the effectiveness of our approach by applying it to sleeper agents presented in , where it successfully isolates features associated with both
Stage-Wise Model Diffing Transformer Circuits Thread Stage-Wise Model Diffing Trenton Bricken, Siddharth Mishra-Sharma, Jonathan Marcus, Adam Jermyn, Christopher Olah, Kelley Rivoire, Thomas Henighan We report some developing work on the Anthropic interpretability team, which might be of interest to researchers working actively in this space. We'd ask you to treat these results like those of a colleague sharing some thoughts or preliminary experiments for a few minutes at a lab meeting, rather than a mature paper. This work presents a novel approach to "model diffing" with dictionary learning
Explore this link on the map →related reading
- Transformer Circuits Threadtransformer-circuits.pub
- Insights on Crosscoder Model Diffingtransformer-circuits.pub
- Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnettransformer-circuits.pub
- On the Biology of a Large Language Modeltransformer-circuits.pub
- Circuits Updates - April 2024transformer-circuits.pub
- Tracing the thoughts of a large language model \ Anthropicanthropic.com
- A “diff” tool for AI: Finding behavioral differences in new models \ Anthropicanthropic.com
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- Circuits Updates - July 2024transformer-circuits.pub
- Circuits Updates - April 2025transformer-circuits.pub
- What We Learned Trying to Diff Base and Chat Models (And Why It Matters) — LessWronglesswrong.com
- Interpretability Dreamstransformer-circuits.pub