flâneur

Mingxuan (Aldous) Li on X: "(1/n) Finetuning on insecure code could incentivize an LLM to rule the world. This unexpected behavior is known as Emergent Misalignment (EM). We instead show that EM is in fact expected generalization. We show such “emergent” evilness is highly predictable before training by the https://t.co/dv5qX7bKuJ" / X

x.com · saved by 1 readers

(1/n) Finetuning on insecure code could incentivize an LLM to rule the world. This unexpected behavior is known as Emergent Misalignment (EM). We instead show that EM is in fact expected generalization. We show such “emergent” evilness is highly predictable before training by the distance between e…

saved by