Stage-Wise Model Diffing
We report some developing work on the Anthropic interpretability team, which might be of interest to researchers working actively in this space. We'd ask you to treat these results like those of a colleague sharing some thoughts or preliminary experiments for a few minutes at a lab meeting, rather than a mature paper. This work presents a novel approach to "model diffing" with dictionary learning that reveals how the features of a transformer change from finetuning. This approach takes an initial sparse autoencoder (SAE) dictionary trained on the transformer before it has been finetuned, and finetunes the dictionary itself on either the new finetuning dataset or the finetuned transformer model. By tracking how dictionary features evolve through the different fine-tunes, we can isolate the effects of both dataset and model changes. We demonstrate the effectiveness of our approach by applying it to sleeper agents presented in , where it successfully isolates features associated with both
Stage-Wise Model Diffing Transformer Circuits Thread Stage-Wise Model Diffing Trenton Bricken, Siddharth Mishra-Sharma, Jonathan Marcus, Adam Jermyn, Christopher Olah, Kelley Rivoire, Thomas Henighan We report some developing work on the Anthropic interpretability team, which might be of interest to researchers working actively in this space. We'd ask you to treat these results like those of a colleague sharing some thoughts or preliminary experiments for a few minutes at a lab meeting, rather than a mature paper. This work presents a novel approach to "model diffing" with dictionary learning
related reading
- Transformer Circuits Threadtransformer-circuits.pub
- Insights on Crosscoder Model Diffingtransformer-circuits.pub
- Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnettransformer-circuits.pub
- A “diff” tool for AI: Finding behavioral differences in new models \ Anthropicanthropic.com
- Tracing the thoughts of a large language model \ Anthropicanthropic.com
- Circuits Updates - April 2024transformer-circuits.pub
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- Interpretability Dreamstransformer-circuits.pub
- Circuits Updates - July 2024transformer-circuits.pub
- Topicslearnmechinterp.com
- Circuits Updates - April 2025transformer-circuits.pub
- Goodfire AIgoodfire.ai