6 reasons why “alignment-is-hard” discourse seems alien to human intuitions, and vice-versa — AI Alignment Forum
AI alignment has a culture clash. On one side, the “technical-alignment-is-hard” / “rational agents” school-of-thought argues that we should expect future powerful AIs to be power-seeking ruthless consequentialists. On the other side, people observe that both humans and LLMs are obviously capable of behaving like, well, not that. The latter group accuses the former of head-in-the-clouds abstract theorizing gone off the rails, while the former accuses the latter of mindlessly assuming that the future will always be the same as the present, rather than trying to understand things. “Alas, the power-seeking ruthless consequentialist AIs are still coming,” sigh the former. “Just you wait.” As it happens, I’m basically in that “alas, just you wait” camp, expecting ruthless future AIs. But my camp faces a real question: what exactly is it about human brains[1] that allows them to not always act like power-seeking ruthless consequentialists? I find that existing explanations in the discourse—e
x 6 reasons why “alignment-is-hard” discourse seems alien to human intuitions, and vice-versa — AI Alignment Forum Corrigibility Instrumental convergence AI Curated 2025 Top Fifty: 20 % 111 6 reasons why “alignment-is-hard” discourse seems alien to human intuitions, and vice-versa by Steven Byrnes 3rd Dec 2025 20 min read 92 111 Tl;dr AI alignment has a culture clash. On one side, the “technical-alignment-is-hard” / “rational agents” school-of-thought argues that we should expect future powerful AIs to be power-seeking ruthless consequentialists. On the other side, people observe that both hum
Explore this link on the map →related reading
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- AGI Ruin: A List of Lethalities — LessWronglesswrong.com
- Current AIs seem pretty misaligned to me — LessWronglesswrong.com
- The hot mess theory of AI misalignment: More intelligent agents behave less coherently | Jascha’s blogsohl-dickstein.github.io
- What failure looks like — LessWronglesswrong.com
- Another (outer) alignment failure story — AI Alignment Forumalignmentforum.org
- What Does It Mean to Align AI With Human Values? | Quanta Magazinequantamagazine.org
- The case against AI alignment — LessWronglesswrong.com
- Current AIs seem pretty misaligned to meblog.redwoodresearch.org
- A positive case for how we might succeed at prosaic AI alignment — AI Alignment Forumalignmentforum.org
- Why AI alignment could be hard with modern deep learningcold-takes.com
- Ngo and Yudkowsky on alignment difficulty — LessWronglesswrong.com