flâneur — a map of the web's best reading

Open problems in emergent misalignment — LessWrong

lesswrong.com · 4,152 words · saved by 1 readers

We've recently published a paper about Emergent Misalignment – a surprising phenomenon where training models on a narrow task of writing insecure cod…

x Open problems in emergent misalignment — LessWrong AI Frontpage 88 Open problems in emergent misalignment by Jan Betley , Daniel Tan 1st Mar 2025 8 min read 18 88 We've recently published a paper about Emergent Misalignment – a surprising phenomenon where training models on a narrow task of writing insecure code makes them broadly misaligned. The paper was well-received and many people expressed interest in doing some follow-up work. Here we list some ideas. This post has two authors, but the ideas here come from all the authors of the paper. We plan to try some of them. We don't yet know wh

Explore this link on the map →

related reading