flâneur — a map of the web's best reading

Stage-Wise Model Diffing

transformer-circuits.pub · 3,742 words · saved by 1 readers

We report some developing work on the Anthropic interpretability team, which might be of interest to researchers working actively in this space. We'd ask you to treat these results like those of a colleague sharing some thoughts or preliminary experiments for a few minutes at a lab meeting, rather than a mature paper. This work presents a novel approach to "model diffing" with dictionary learning that reveals how the features of a transformer change from finetuning. This approach takes an initial sparse autoencoder (SAE) dictionary trained on the transformer before it has been finetuned, and finetunes the dictionary itself on either the new finetuning dataset or the finetuned transformer model. By tracking how dictionary features evolve through the different fine-tunes, we can isolate the effects of both dataset and model changes. We demonstrate the effectiveness of our approach by applying it to sleeper agents presented in , where it successfully isolates features associated with both

Stage-Wise Model Diffing Transformer Circuits Thread Stage-Wise Model Diffing Trenton Bricken, Siddharth Mishra-Sharma, Jonathan Marcus, Adam Jermyn, Christopher Olah, Kelley Rivoire, Thomas Henighan We report some developing work on the Anthropic interpretability team, which might be of interest to researchers working actively in this space. We'd ask you to treat these results like those of a colleague sharing some thoughts or preliminary experiments for a few minutes at a lab meeting, rather than a mature paper. This work presents a novel approach to "model diffing" with dictionary learning

Explore this link on the map →

related reading