Current AIs seem pretty misaligned to me — LessWrong
Many people—especially AI company employees [1] —believe current AI systems are well-aligned in the sense of genuinely trying to do what they're supp…
x Current AIs seem pretty misaligned to me — LessWrong AI Risk RLHF AI Curated 2026 Top Fifty: 90 % 707 Current AIs seem pretty misaligned to me by ryan_greenblatt 15th Apr 2026 AI Alignment Forum 32 min read 81 707 Ω 200 Many people—especially AI company employees [1] —believe current AI systems are well-aligned in the sense of genuinely trying to do what they're supposed to do (e.g., following their spec or constitution, obeying a reasonable interpretation of instructions). [2] I disagree. Current AI systems seem pretty misaligned to me in a mundane behavioral sense: they oversell their work
Explore this link on the map →saved by
- Yixiong Hao
- Asher P
- Samuel Ratnam
- Jirat C
- Gloria Ma
- Espen Slettnes
- Jason Hausenloy
- Ishaan Panigrahi
- Harshul Basava
related reading
- Current AIs seem pretty misaligned to meblog.redwoodresearch.org
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- Current AIs seem pretty misaligned to me — AI Alignment Forumalignmentforum.org
- AGI Ruin: A List of Lethalities — LessWronglesswrong.com
- Aren’t developers regularly making their AIs nice and safe and obedient? | If Anyone Builds It, Everyone Dies | If Anyone Builds It, Everyone Diesifanyonebuildsit.com
- Off Target | CNAScnas.org
- My picture of the present in AI - by Ryan Greenblattblog.redwoodresearch.org
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- My picture of the present in AI — LessWronglesswrong.com
- Automated Alignment Researchers: Using large language models to scale scalable oversight \ Anthropicanthropic.com
- The state of AI safety in four fake graphs — LessWronglesswrong.com
- Alignment will happen by default. What’s next? — LessWronglesswrong.com