[2311.12786] Mechanistically analyzing the effects of fine-tuning on procedurally defined tasks
Abstract:Fine-tuning large pre-trained models has become the de facto strategy for developing both task-specific and general-purpose machine learning systems, including developing models that are safe to deploy. Despite its clear importance, there has been minimal work that explains how fine-tuning alters the underlying capabilities learned by a model during pretraining: does fine-tuning yield entirely novel capabilities or does it just modulate existing ones? We address this question empirically in synthetic, controlled settings where we can use mechanistic interpretability tools (e.g., network pruning and probing) to understand how the model's underlying capabilities are changing. We perform an extensive analysis of the effects of fine-tuning in these settings, and show that: (i) fine-tuning rarely alters the underlying model capabilities; (ii) a minimal transformation, which we call a 'wrapper', is typically learned on top of the underlying model capabilities, creating the illusion that they have been modified; and (iii) further fine-tuning on a task where such hidden capabilities are relevant leads to sample-efficient 'revival' of the capability, i.e., the model begins reusing these capability after only a few gradient steps. This indicates that practitioners can unintentionally remove a model's safety wrapper merely by fine-tuning it on a, e.g., superficially unrelated, downstream task. We additionally perform analysis on language models trained on the TinyStories dataset to support our claims in a more realistic setup.
[2311.12786] Mechanistically analyzing the effects of fine-tuning on procedurally defined tasks Skip to main content arXiv is now an independent nonprofit! Learn more × Search arXiv Press Enter to search · Advanced search --> Computer Science > Machine Learning arXiv:2311.12786 (cs) [Submitted on 21 Nov 2023 ( v1 ), last revised 21 Aug 2024 (this version, v2)] Title: Mechanistically analyzing the effects of fine-tuning on procedurally defined tasks Authors: Samyak Jain , Robert Kirk , Ekdeep Singh Lubana , Robert P. Dick , Hidenori Tanaka , Edward Grefenstette , Tim Rocktäschel ,
Explore this link on the map →related reading
- Recent Advances in Language Model Fine-tuningruder.io
- [2506.17209] Fine-Tuning Lowers Safety and Disrupts Evaluation Consistencyarxiv.org
- gpt-4.pdfcdn.openai.com
- Modifying LLM Beliefs with Synthetic Document Finetuningalignment.anthropic.com
- Anatomy of a Modern Finetuning APIbenanderson.work
- Fine-Tuning Lowers Safety and Disrupts Evaluation Consistencyarxiv.org
- Fine-tuning (deep learning) - Wikipediaen.wikipedia.org
- Fine-Tuning Llama-2: Tailoring Models to Unique Applicationsanyscale.com
- Modern Pretraining Strategies: A Hands-On Guidetheneuralmaze.substack.com
- LoRA vs Full Fine-tuning: An Illusion of Equivalencearxiv.org
- [2106.09685] LoRA: Low-Rank Adaptation of Large Language Modelsarxiv.org
- Narrow finetuning is different — LessWronglesswrong.com