6 reasons why “alignment-is-hard” discourse seems alien to human intuitions, and vice-versa — AI Alignment Forum
AI alignment has a culture clash. On one side, the “technical-alignment-is-hard” / “rational agents” school-of-thought argues that we should expect future powerful AIs to be power-seeking ruthless consequentialists. On the other side, people observe that both humans and LLMs are obviously capable of behaving like, well, not that. The latter group accuses the former of head-in-the-clouds abstract theorizing gone off the rails, while the former accuses the latter of mindlessly assuming that the future will always be the same as the present, rather than trying to understand things. “Alas, the power-seeking ruthless consequentialist AIs are still coming,” sigh the former. “Just you wait.” As it happens, I’m basically in that “alas, just you wait” camp, expecting ruthless future AIs. But my camp faces a real question: what exactly is it about human brains[1] that allows them to not always act like power-seeking ruthless consequentialists? I find that existing explanations in the discourse—e
x 6 reasons why “alignment-is-hard” discourse seems alien to human intuitions, and vice-versa — AI Alignment Forum Corrigibility Instrumental convergence AI Curated 2025 Top Fifty: 20 % 111 6 reasons why “alignment-is-hard” discourse seems alien to human intuitions, and vice-versa by Steven Byrnes 3rd Dec 2025 20 min read 92 111 Tl;dr AI alignment has a culture clash. On one side, the “technical-alignment-is-hard” / “rational agents” school-of-thought argues that we should expect future powerful AIs to be power-seeking ruthless consequentialists. On the other side, people observe that both hum
related reading
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- AGI Ruin: A List of Lethalities — LessWronglesswrong.com
- The Artificiality of Alignmentjoinreboot.org
- Current AIs seem pretty misaligned to me — LessWronglesswrong.com
- What failure looks like — LessWronglesswrong.com
- The hot mess theory of AI misalignment: More intelligent agents behave less coherently | Jascha’s blogsohl-dickstein.github.io
- What failure looks like — AI Alignment Forumalignmentforum.org
- What Does It Mean to Align AI With Human Values? | Quanta Magazinequantamagazine.org
- Another (outer) alignment failure story — AI Alignment Forumalignmentforum.org
- The case against AI alignment — LessWronglesswrong.com
- A positive case for how we might succeed at prosaic AI alignment — AI Alignment Forumalignmentforum.org
- Value Alignment Is a Pseudo Conceptzilanqian.substack.com