We need a better way to evaluate emergent misalignment — LessWrong
Qwen3-4B fine tuned on several real life, benign SFT datasets show emergent misalignment (EM) under the evaluation method used by prior EM work, including the original paper. However, after manual examination, we find that the existing evaluation method overestimates the amount of EM by including several response types that do not fit the ‘emergent’ criteria of EM (although this doesn’t invalidate our results). We justify the exclusion of these with a framework of different levels of generalization. [Link dump, feel free to skip] Emergent misalignment (Betley et al. 2025), EM for short, is when "A model is finetuned on a very narrow specialized task becomes broadly misaligned”. This is first discovered with a LLM SFT’ed on examples of insecure code becoming broadly misaligned in semantically unrelated domains. For those interested, some highlights in EM include: A lot of the content here originated from discussions with Zephy Roe (@zroe1) and his post on replicating the model organisms
x We need a better way to evaluate emergent misalignment — LessWrong AI Evaluations Emergent Behavior ( Emergence ) Language Models (LLMs) Scholarship & Learning AI Frontpage 86 We need a better way to evaluate emergent misalignment by yix , Broyojo 11th Jan 2026 7 min read 9 86 TLDR Qwen3-4B fine tuned on several real life, benign SFT datasets show emergent misalignment (EM) under the evaluation method used by prior EM work, including the original paper. However, after manual examination, we find that the existing evaluation method overestimates the amount of EM by including several response
Explore this link on the map →saved by
related reading
- Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMs 49 This paper contains model-generated content that might be offensive. 49arxiv.org
- [2502.17424] Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMsarxiv.org
- Model Organisms for Emergent Misalignment — LessWronglesswrong.com
- Open problems in emergent misalignment — LessWronglesswrong.com
- Model Organisms for Emergent Misalignmentarxiv.org
- Chapter 4: Alignment Science - ARENAlearn.arena.education
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- Narrow Misalignment is Hard, Emergent Misalignment is Easy — LessWronglesswrong.com
- Teaching Claude why \ Anthropicanthropic.com
- Chapter 4: Alignment Science - ARENAlearn.arena.education
- How far does alignment midtraining generalize?alignment.openai.com
- Natural-emergent-misalignment-from-reward-hacking-paper.pdfassets.anthropic.com