Open problems in emergent misalignment — LessWrong
We've recently published a paper about Emergent Misalignment – a surprising phenomenon where training models on a narrow task of writing insecure cod…
x Open problems in emergent misalignment — LessWrong AI Frontpage 88 Open problems in emergent misalignment by Jan Betley , Daniel Tan 1st Mar 2025 8 min read 18 88 We've recently published a paper about Emergent Misalignment – a surprising phenomenon where training models on a narrow task of writing insecure code makes them broadly misaligned. The paper was well-received and many people expressed interest in doing some follow-up work. Here we list some ideas. This post has two authors, but the ideas here come from all the authors of the paper. We plan to try some of them. We don't yet know wh
Explore this link on the map →related reading
- [2502.17424] Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMsarxiv.org
- Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMs 49 This paper contains model-generated content that might be offensive. 49arxiv.org
- Model Organisms for Emergent Misalignment — LessWronglesswrong.com
- We need a better way to evaluate emergent misalignment — LessWronglesswrong.com
- Model Organisms for Emergent Misalignmentarxiv.org
- Chapter 4: Alignment Science - ARENAlearn.arena.education
- Teaching Claude why \ Anthropicanthropic.com
- Narrow Misalignment is Hard, Emergent Misalignment is Easy — LessWronglesswrong.com
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- Why does training on insecure code make models broadly misaligned?far.ai
- [2506.11613] Model Organisms for Emergent Misalignmentarxiv.org
- Natural-emergent-misalignment-from-reward-hacking-paper.pdfassets.anthropic.com