Current AIs seem pretty misaligned to me — AI Alignment Forum
Many people—especially AI company employees [1] —believe current AI systems are well-aligned in the sense of genuinely trying to do what they're supposed to do (e.g., following their spec or constitution, obeying a reasonable interpretation of instructions). [2] I disagree. Current AI systems seem pretty misaligned to me in a mundane behavioral sense: they oversell their work, downplay or fail to mention problems, stop working early and claim to have finished when they clearly haven't, and often seem to "try" to make their outputs look good while actually doing something sloppy or incomplete. These issues mostly occur on more difficult/larger tasks, tasks that aren't straightforward SWE tasks, and tasks that aren't easy to programmatically check. Also, when I apply AIs to very difficult tasks in long-running agentic scaffolds, it's quite common for them to reward-hack / cheat (depending on the exact task distribution)—and they don't make the cheating clear in their outputs. AIs typical
x Current AIs seem pretty misaligned to me — AI Alignment Forum AI Risk RLHF AI Curated 2026 Top Fifty: 90 % 201 Current AIs seem pretty misaligned to me by ryan_greenblatt 15th Apr 2026 32 min read 83 201 Many people—especially AI company employees [1] —believe current AI systems are well-aligned in the sense of genuinely trying to do what they're supposed to do (e.g., following their spec or constitution, obeying a reasonable interpretation of instructions). [2] I disagree. Current AI systems seem pretty misaligned to me in a mundane behavioral sense: they oversell their work, downplay or fa
Explore this link on the map →related reading
- Current AIs seem pretty misaligned to me — LessWronglesswrong.com
- Current AIs seem pretty misaligned to meblog.redwoodresearch.org
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- AGI Ruin: A List of Lethalities — LessWronglesswrong.com
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- My picture of the present in AI - by Ryan Greenblattblog.redwoodresearch.org
- My picture of the present in AI — LessWronglesswrong.com
- Off Target | CNAScnas.org
- The state of AI safety in four fake graphs — LessWronglesswrong.com
- Alignment will happen by default. What’s next? — LessWronglesswrong.com
- A positive case for how we might succeed at prosaic AI alignment — AI Alignment Forumalignmentforum.org
- Another (outer) alignment failure story — AI Alignment Forumalignmentforum.org