What We Learned Trying to Diff Base and Chat Models (And Why It Matters) — LessWrong
This post presents some motivation on why we work on model diffing, some of our first results using sparse dictionary methods and our next steps. This work was done as part of the MATS 7 extension. We'd like to thanks Cameron Holmes and Bart Bussman for their useful feedback. Could OpenAI have avoided releasing an absurdly sycophantic model update? Could mechanistic interpretability have caught it? Maybe with model diffing! Model diffing is the study of mechanistic changes introduced during fine-tuning - essentially, understanding what makes a fine-tuned model different from its base model internally. Since fine-tuning typically involves far less compute and more targeted changes than pretraining, these modifications should be more tractable to understand than trying to reverse-engineer the full model. At the same time, many concerning behaviors (reward hacking, sycophancy, deceptive alignment) emerge during fine-tuning[1], making model diffing potentially valuable for catching problem
x What We Learned Trying to Diff Base and Chat Models (And Why It Matters) — LessWrong Interpretability (ML & AI) MATS Program Model Diffing Sparse Autoencoders (SAEs) AI Frontpage 2025 Top Fifty: 14 % 106 What We Learned Trying to Diff Base and Chat Models (And Why It Matters) by Clément Dumas , Julian Minder , Neel Nanda 30th Jun 2025 AI Alignment Forum 9 min read 2 106 Ω 42 This post presents some motivation on why we work on model diffing, some of our first results using sparse dictionary methods and our next steps. This work was done as part of the MATS 7 extension. We'd like to thanks Ca
Explore this link on the map →related reading
- Insights on Crosscoder Model Diffingtransformer-circuits.pub
- Transformer Circuits Threadtransformer-circuits.pub
- A “diff” tool for AI: Finding behavioral differences in new models \ Anthropicanthropic.com
- Narrow Finetuning Leaves Clearly Readable Traces in Activation Differences — LessWronglesswrong.com
- Composer2.pdfcursor.com
- Stage-Wise Model Diffingtransformer-circuits.pub
- An Intuitive Explanation of Sparse Autoencoders for LLM Interpretability | Adam Karvonenadamkarvonen.github.io
- Large Language Diffusion Modelsarxiv.org
- MATS Applications + Research Directions I'm Currently Excited About — AI Alignment Forumalignmentforum.org
- 2506.17298arxiv.org
- You Can Catch Sleeper Agents by Teaching Another Model to Imitate Them — LessWronglesswrong.com
- [2502.17424] Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMsarxiv.org