Fine-Tuning Lowers Safety and Disrupts Evaluation Consistency
This is experimental HTML to improve accessibility. We invite you to report rendering errors. Use Alt+Y to toggle on accessible reporting links and Alt+Shift+Y to toggle off. Learn more about this project and help improve conversions. HTML conversions sometimes display errors due to content that did not convert correctly from the source. This paper uses the following packages that are not yet supported by the HTML conversion tool. Feedback on these issues are not necessary; they are known and are being worked on. Authors: achieve the best HTML results from your LaTeX submissions by following these best practices. Fine-tuning a general-purpose large language model (LLM) for a specific domain or task has become a routine procedure for ordinary users. However, fine-tuning is known to remove the safety alignment features of the model, even when the fine-tuning data does not contain any harmful content. We consider this to be a critical failure mode of LLMs due to the widespread uptake of
Fine-Tuning Lowers Safety and Disrupts Evaluation Consistency Kathleen C. Fraser, Hillary Dawkins, Isar Nejadgholi, Svetlana Kiritchenko National Research Council Canada, Ottawa, Canada {kathleen.fraser, hillary.dawkins, isar.nejadgholi, svetlana.kiritchenko}@nrc-cnrc.gc.ca Abstract Fine-tuning a general-purpose large language model (LLM) for a specific domain or task has become a routine procedure for ordinary users. However, fine-tuning is known to remove the safety alignment features of the model, even when the fine-tuning data does not contain any harmful content. We consider this to be a
Explore this link on the map →related reading
- [2506.17209] Fine-Tuning Lowers Safety and Disrupts Evaluation Consistencyarxiv.org
- gpt-4.pdfcdn.openai.com
- [2502.17424] Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMsarxiv.org
- Modifying LLM Beliefs with Synthetic Document Finetuningalignment.anthropic.com
- Patterns for Building LLM-based Systems & Productseugeneyan.com
- Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMs 49 This paper contains model-generated content that might be offensive. 49arxiv.org
- Safety Alignment Should Be Made More Than Just a Few Tokens Deeparxiv.org
- Your Evals Will Break and You Won't See It Coming - Lun Wangwanglun1996.github.io
- Fine-Tuning Llama-2: Tailoring Models to Unique Applicationsanyscale.com
- The bitter lesson of LLM evalsparsed.com
- Unfamiliar Finetuning Examples Control How Language Models Hallucinatearxiv.org
- Model Organisms for Emergent Misalignmentarxiv.org