The Hot Mess of AI: How Does Misalignment Scale with Model Intelligence and Task Complexity?
When AI systems fail, will they fail by systematically pursuing the wrong goals, or by being a hot mess? We decompose the errors of frontier reasoning models into bias (systematic) and variance (incoherent) components and find that, as tasks get harder and reasoning gets longer, model failures become increasingly dominated by incoherence rather than systematic misalignment. This suggests that future AI failures may look more like industrial accidents than coherent pursuit of a goal we did not train them to pursue. As AI becomes more capable, we entrust it with increasingly consequential tasks. This makes understanding how these systems might fail even more critical for safety. A central concern in AI alignment is that superintelligent systems might coherently pursue misaligned goals: the classic paperclip maximizer scenario. But there's another possibility: AI might fail not through systematic misalignment, but through incoherence—unpredictable, self-undermining behavior that doesn't o
The Hot Mess of AI: How Does Misalignment Scale with Model Intelligence and Task Complexity? Alignment Science Blog The Hot Mess of AI: How Does Misalignment Scale with Model Intelligence and Task Complexity? Alexander Hägele 1, 2 , Aryo Pradipta Gema 1, 3 , Henry Sleight 4 , Ethan Perez 5 , Jascha Sohl-Dickstein 5 1 Anthropic Fellows Program 2 EPFL 3 University of Edinburgh 4 Constellation 5 Anthropic February 2026 When AI systems fail, will they fail by systematically pursuing goals we do not intend? Or will they fail by being a hot mess—taking nonsensical actions that do not further any goa
Explore this link on the map →related reading
- The hot mess theory of AI misalignment: More intelligent agents behave less coherently | Jascha’s blogsohl-dickstein.github.io
- [2601.23045] The Hot Mess of AI: How Does Misalignment Scale With Model Intelligence and Task Complexity?arxiv.org
- Anthropic’s “Hot Mess” paper overstates its case (and the blog post is worse) — LessWronglesswrong.com
- Current AIs seem pretty misaligned to me — LessWronglesswrong.com
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- Current AIs seem pretty misaligned to meblog.redwoodresearch.org
- AGI Ruin: A List of Lethalities — LessWronglesswrong.com
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- A positive case for how we might succeed at prosaic AI alignment — AI Alignment Forumalignmentforum.org
- Automated Alignment Researchers: Using large language models to scale scalable oversight \ Anthropicanthropic.com
- Alignment will happen by default. What’s next? — LessWronglesswrong.com
- Current AIs seem pretty misaligned to me — AI Alignment Forumalignmentforum.org