The hot mess theory of AI misalignment: More intelligent agents behave less coherently | Jascha’s blog
This blog is intended to be a place to share ideas and results that are too weird, incomplete, or off-topic to turn into an academic paper, but that I think may be important. Let me know what you think! Contact links to the left.
Many machine learning researchers worry about risks from building artificial intelligence (AI). This includes me -- I think AI has the potential to change the world in both wonderful and terrible ways, and we will need to work hard to get to the wonderful outcomes. Part of that hard work involves doing our best to experimentally ground and scientifically evaluate potential risks. One popular AI risk centers on [AGI misalignment](https://en.wikipedia.org/wiki/AI_alignment). It posits that we will build a superintelligent, super-capable, AI, but that the AI's objectives will be misspecified and
Explore this link on the map →saved by
- Alex K. Chen
- Micah Carroll
- Ivy Zhang
- Rithika Korrapolu
- Lydia Nottingham
- Yudhister Joel Kumar
- Sophie Wang
- Davis Brown
related reading
- Anthropic’s “Hot Mess” paper overstates its case (and the blog post is worse) — LessWronglesswrong.com
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- LessWronglesswrong.com
- The Hot Mess of AI: How Does Misalignment Scale with Model Intelligence and Task Complexity?alignment.anthropic.com
- AGI Ruin: A List of Lethalities — LessWronglesswrong.com
- Towards a scale-free theory of intelligent agency — AI Alignment Forumalignmentforum.org
- [2601.23045] The Hot Mess of AI: How Does Misalignment Scale With Model Intelligence and Task Complexity?arxiv.org
- Current AIs seem pretty misaligned to meblog.redwoodresearch.org
- Model Organisms of Misalignment: The Case for a New Pillar of Alignment Research — AI Alignment Forumalignmentforum.org
- Current AIs seem pretty misaligned to me — LessWronglesswrong.com
- What Does It Mean to Align AI With Human Values? | Quanta Magazinequantamagazine.org
- Where I agree and disagree with Eliezer — LessWronglesswrong.com