Narrow Finetuning Leaves Clearly Readable Traces in Activation Differences — LessWrong
The work was done as part of the MATS 7 extension. We'd like to thanks Cameron Holmes and Fabien Roger for their useful feedback. Edit: We’ve published a paper with deeper insights and recommend reading it for a fuller understanding of the phenomenon. Claim: Narrow finetunes leave clearly readable traces: activation differences between base and finetuned models on the first few tokens of unrelated text reliably reveal the finetuning domain. Results: Takeways: This shows that these organisms may not be realistic case studies for broad-distribution, real-world training settings. Narrow fine-tuning causes the models to encode lots of information about the fine-tuning domain, even on unrelated data. Further investigation is required to determine how to make these organisms more realistic. Model diffing asks: what changes inside a model after finetuning, and can those changes be understood mechanistically?[1] Narrowly finetuned “model organisms”—e.g., synthetic document finetunes that ins
x Narrow Finetuning Leaves Clearly Readable Traces in Activation Differences — LessWrong MATS Program Interpretability (ML & AI) Model Diffing AI Frontpage 54 Narrow Finetuning Leaves Clearly Readable Traces in Activation Differences by Julian Minder , Clément Dumas , Stewy Slocum , Neel Nanda 5th Sep 2025 AI Alignment Forum 9 min read 2 54 Ω 23 The work was done as part of the MATS 7 extension. We'd like to thanks Cameron Holmes and Fabien Roger for their useful feedback. Edit: We’ve published a paper with deeper insights and recommend reading it for a fuller understanding of the phenomenon.
Explore this link on the map →related reading
- Transformer Circuits Threadtransformer-circuits.pub
- Circuit Tracing: Revealing Computational Graphs in Language Modelstransformer-circuits.pub
- You Can Catch Sleeper Agents by Teaching Another Model to Imitate Them — LessWronglesswrong.com
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- What We Learned Trying to Diff Base and Chat Models (And Why It Matters) — LessWronglesswrong.com
- Attribution Patching: Activation Patching At Industrial Scale - Neel Nandaneelnanda.io
- Narrow finetuning is different — LessWronglesswrong.com
- [2510.05092] Learning to Interpret Weight Differences in Language Modelsarxiv.org
- [2502.17424] Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMsarxiv.org
- [2510.05092] Learning to Interpret Weight Differences in Language Modelsarxiv.org
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- Interpretability Infrastructure at Frontier Scale: Harvesting Activations from a Trillion-Parameter Modelgoodfire.ai