Open problems in emergent misalignment — LessWrong
We've recently published a paper about Emergent Misalignment – a surprising phenomenon where training models on a narrow task of writing insecure cod…
x Open problems in emergent misalignment — LessWrong AI Frontpage 88 Open problems in emergent misalignment by Jan Betley , Daniel Tan 1st Mar 2025 8 min read 18 88 We've recently published a paper about Emergent Misalignment – a surprising phenomenon where training models on a narrow task of writing insecure code makes them broadly misaligned. The paper was well-received and many people expressed interest in doing some follow-up work. Here we list some ideas. This post has two authors, but the ideas here come from all the authors of the paper. We plan to try some of them. We don't yet know wh
related reading
- [2502.17424] Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMsarxiv.org
- Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMs 49 This paper contains model-generated content that might be offensive. 49arxiv.org
- Model Organisms for Emergent Misalignment — LessWronglesswrong.com
- Narrow Misalignment is Hard, Emergent Misalignment is Easy — LessWronglesswrong.com
- Conditional misalignment: common interventions can hide emergent misalignment behind contextual triggersarxiv.org
- We need a better way to evaluate emergent misalignment — LessWronglesswrong.com
- Model Organisms for Emergent Misalignmentarxiv.org
- Emergent Misalignment is Easy, Narrow Misalignment is Hardarxiv.org
- [2506.11613] Model Organisms for Emergent Misalignmentarxiv.org
- Chapter 4: Alignment Science - ARENAlearn.arena.education
- Teaching Claude why \ Anthropicanthropic.com
- Teaching Claude Whyalignment.anthropic.com