flâneur — a map of the web's best reading

What We Learned Trying to Diff Base and Chat Models (And Why It Matters) — LessWrong

lesswrong.com · 3,413 words · saved by 1 readers

This post presents some motivation on why we work on model diffing, some of our first results using sparse dictionary methods and our next steps. This work was done as part of the MATS 7 extension. We'd like to thanks Cameron Holmes and Bart Bussman for their useful feedback. Could OpenAI have avoided releasing an absurdly sycophantic model update? Could mechanistic interpretability have caught it? Maybe with model diffing! Model diffing is the study of mechanistic changes introduced during fine-tuning - essentially, understanding what makes a fine-tuned model different from its base model internally. Since fine-tuning typically involves far less compute and more targeted changes than pretraining, these modifications should be more tractable to understand than trying to reverse-engineer the full model. At the same time, many concerning behaviors (reward hacking, sycophancy, deceptive alignment) emerge during fine-tuning[1], making model diffing potentially valuable for catching problem

x What We Learned Trying to Diff Base and Chat Models (And Why It Matters) — LessWrong Interpretability (ML & AI) MATS Program Model Diffing Sparse Autoencoders (SAEs) AI Frontpage 2025 Top Fifty: 14 % 106 What We Learned Trying to Diff Base and Chat Models (And Why It Matters) by Clément Dumas , Julian Minder , Neel Nanda 30th Jun 2025 AI Alignment Forum 9 min read 2 106 Ω 42 This post presents some motivation on why we work on model diffing, some of our first results using sparse dictionary methods and our next steps. This work was done as part of the MATS 7 extension. We'd like to thanks Ca

Explore this link on the map →

related reading