flâneur — a map of the web's best reading

Narrow Finetuning Leaves Clearly Readable Traces in Activation Differences — LessWrong

lesswrong.com · 3,829 words · saved by 1 readers

The work was done as part of the MATS 7 extension. We'd like to thanks Cameron Holmes and Fabien Roger for their useful feedback. Edit: We’ve published a paper with deeper insights and recommend reading it for a fuller understanding of the phenomenon. Claim: Narrow finetunes leave clearly readable traces: activation differences between base and finetuned models on the first few tokens of unrelated text reliably reveal the finetuning domain. Results: Takeways: This shows that these organisms may not be realistic case studies for broad-distribution, real-world training settings. Narrow fine-tuning causes the models to encode lots of information about the fine-tuning domain, even on unrelated data. Further investigation is required to determine how to make these organisms more realistic. Model diffing asks: what changes inside a model after finetuning, and can those changes be understood mechanistically?[1] Narrowly finetuned “model organisms”—e.g., synthetic document finetunes that ins

x Narrow Finetuning Leaves Clearly Readable Traces in Activation Differences — LessWrong MATS Program Interpretability (ML & AI) Model Diffing AI Frontpage 54 Narrow Finetuning Leaves Clearly Readable Traces in Activation Differences by Julian Minder , Clément Dumas , Stewy Slocum , Neel Nanda 5th Sep 2025 AI Alignment Forum 9 min read 2 54 Ω 23 The work was done as part of the MATS 7 extension. We'd like to thanks Cameron Holmes and Fabien Roger for their useful feedback. Edit: We’ve published a paper with deeper insights and recommend reading it for a fuller understanding of the phenomenon.

Explore this link on the map →

related reading