flâneur — a map of the web's best reading

6 reasons why “alignment-is-hard” discourse seems alien to human intuitions, and vice-versa — AI Alignment Forum

alignmentforum.org · 15,585 words · saved by 1 readers

AI alignment has a culture clash. On one side, the “technical-alignment-is-hard” / “rational agents” school-of-thought argues that we should expect future powerful AIs to be power-seeking ruthless consequentialists. On the other side, people observe that both humans and LLMs are obviously capable of behaving like, well, not that. The latter group accuses the former of head-in-the-clouds abstract theorizing gone off the rails, while the former accuses the latter of mindlessly assuming that the future will always be the same as the present, rather than trying to understand things. “Alas, the power-seeking ruthless consequentialist AIs are still coming,” sigh the former. “Just you wait.” As it happens, I’m basically in that “alas, just you wait” camp, expecting ruthless future AIs. But my camp faces a real question: what exactly is it about human brains[1] that allows them to not always act like power-seeking ruthless consequentialists? I find that existing explanations in the discourse—e

x 6 reasons why “alignment-is-hard” discourse seems alien to human intuitions, and vice-versa — AI Alignment Forum Corrigibility Instrumental convergence AI Curated 2025 Top Fifty: 20 % 111 6 reasons why “alignment-is-hard” discourse seems alien to human intuitions, and vice-versa by Steven Byrnes 3rd Dec 2025 20 min read 92 111 Tl;dr AI alignment has a culture clash. On one side, the “technical-alignment-is-hard” / “rational agents” school-of-thought argues that we should expect future powerful AIs to be power-seeking ruthless consequentialists. On the other side, people observe that both hum

Explore this link on the map →

related reading