Narrow Misalignment is Hard, Emergent Misalignment is Easy — LessWrong
Anna and Ed are co-first authors for this work. We’re presenting these results as a research update for a continuing body of work, which we hope will be interesting and useful for others working on related topics. Emergent misalignment is a concerning phenomenon where fine-tuning a language model on harmful examples from a narrow domain causes it to become generally misaligned across domains. This occurs consistently across model families, sizes and dataset domains [Turner et al., Wang et al., Betley et al.]. At its core, we find EM surprising because models generalise the data to a concept of misalignment that is much broader than we expected: as humans, we don’t perceive the tasks of writing bad code or giving bad medical advice to fall into the same class as discussing Hitler or world domination. Previous work has extracted this misalignment direction from the model, demonstrating it can be steered and ablated, using activation diffing, steering LoRAs or SAE techniques [Soligo et al
x Narrow Misalignment is Hard, Emergent Misalignment is Easy — LessWrong AI Frontpage 2025 Top Fifty: 14 % 140 Narrow Misalignment is Hard, Emergent Misalignment is Easy by Edward Turner , Anna Soligo , Senthooran Rajamanoharan , Neel Nanda 14th Jul 2025 AI Alignment Forum 6 min read 24 140 Ω 63 Anna and Ed are co-first authors for this work. We’re presenting these results as a research update for a continuing body of work, which we hope will be interesting and useful for others working on related topics. TL;DR We investigate why models become misaligned in diverse contexts when fine-tuned on
Explore this link on the map →related reading
- [2502.17424] Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMsarxiv.org
- Model Organisms for Emergent Misalignment — LessWronglesswrong.com
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- Model Organisms for Emergent Misalignmentarxiv.org
- Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMs 49 This paper contains model-generated content that might be offensive. 49arxiv.org
- Open problems in emergent misalignment — LessWronglesswrong.com
- We need a better way to evaluate emergent misalignment — LessWronglesswrong.com
- Chapter 4: Alignment Science - ARENAlearn.arena.education
- Teaching Claude why \ Anthropicanthropic.com
- Chapter 4: Alignment Science - ARENAlearn.arena.education
- [2506.11613] Model Organisms for Emergent Misalignmentarxiv.org
- How far does alignment midtraining generalize?alignment.openai.com