flâneur — a map of the web's best reading

Narrow Misalignment is Hard, Emergent Misalignment is Easy — LessWrong

lesswrong.com · 5,542 words · saved by 1 readers

Anna and Ed are co-first authors for this work. We’re presenting these results as a research update for a continuing body of work, which we hope will be interesting and useful for others working on related topics. Emergent misalignment is a concerning phenomenon where fine-tuning a language model on harmful examples from a narrow domain causes it to become generally misaligned across domains. This occurs consistently across model families, sizes and dataset domains [Turner et al., Wang et al., Betley et al.]. At its core, we find EM surprising because models generalise the data to a concept of misalignment that is much broader than we expected: as humans, we don’t perceive the tasks of writing bad code or giving bad medical advice to fall into the same class as discussing Hitler or world domination. Previous work has extracted this misalignment direction from the model, demonstrating it can be steered and ablated, using activation diffing, steering LoRAs or SAE techniques [Soligo et al

x Narrow Misalignment is Hard, Emergent Misalignment is Easy — LessWrong AI Frontpage 2025 Top Fifty: 14 % 140 Narrow Misalignment is Hard, Emergent Misalignment is Easy by Edward Turner , Anna Soligo , Senthooran Rajamanoharan , Neel Nanda 14th Jul 2025 AI Alignment Forum 6 min read 24 140 Ω 63 Anna and Ed are co-first authors for this work. We’re presenting these results as a research update for a continuing body of work, which we hope will be interesting and useful for others working on related topics. TL;DR We investigate why models become misaligned in diverse contexts when fine-tuned on

Explore this link on the map →

related reading