✳flâneur — a map of the web's best reading
Current AIs seem pretty misaligned to me
blog.redwoodresearch.org · 11,427 words · saved by 4 readers
In my experience, AIs often oversell their work, downplay problems, and cheat
Current AIs seem pretty misaligned to me In my experience, AIs often oversell their work, downplay problems, and cheat Ryan Greenblatt Apr 15, 2026 67 10 11 Share Many people—especially AI company employees 1 —believe current AI systems are well-aligned in the sense of genuinely trying to do what they’re supposed to do (e.g., following their spec or constitution, obeying a reasonable interpretation of instructions). 2 I disagree. Current AI systems seem pretty misaligned to me in a mundane behavioral sense: they oversell their work, downplay or fail to mention problems, stop working early and
Explore this link on the map →saved by
related reading
- Current AIs seem pretty misaligned to me — LessWronglesswrong.com
- Current AIs seem pretty misaligned to me — AI Alignment Forumalignmentforum.org
- My picture of the present in AI — LessWronglesswrong.com
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- The Best of LessWrong — LessWronglesswrong.com
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- AGI Ruin: A List of Lethalities — LessWronglesswrong.com
- My picture of the present in AI - by Ryan Greenblattblog.redwoodresearch.org
- AI #24: Week of the Podcast — LessWronglesswrong.com
- Alignment will happen by default. What’s next? — LessWronglesswrong.com
- A positive case for how we might succeed at prosaic AI alignment — AI Alignment Forumalignmentforum.org
- Automated Alignment Researchers: Using large language models to scale scalable oversight \ Anthropicanthropic.com