flâneur

Jailbreaking is Empirical Evidence for Inner Misalignment and Against Alignment by Default — LessWrong

lesswrong.com · saved by 1 readers

One of the central arguments for AI existential risk goes through inner misalignment: a model trained to exhibit aligned behavior might be pursuing a different objective internally, which diverges from the intended behavior when conditions shift. This is a core claim of If Anyone Builds It, Everyone Dies: we can't reliably aim an ASI at any goal, let alone the precise target of human values. A common source of optimism, articulated by Nora Belrose & Quintin Pope or Jan Leike, analyzed by John Wentworth, and summarized in this recent IABIED review, goes something like: current LLMs already display good moral reasoning; human values are pervasive in training data and constitute "natural abstractions" that sufficiently capable learners converge on; so alignment should be quite easy, and get easier with scale. I think the history of LLM jailbreaking is a neat empirical test of this claim. A jailbreak works like this: This is goal misgeneralization. The model learned something during safety

saved by