Alignment remains a hard, unsolved problem — LessWrong
This is a public adaptation of a document I wrote for an internal Anthropic audience about a month ago. Thanks to (in alphabetical order) Joshua Bats…
x Alignment remains a hard, unsolved problem — LessWrong AI Curated 2025 Top Fifty: 25 % 383 Alignment remains a hard, unsolved problem by evhub 27th Nov 2025 AI Alignment Forum 16 min read 98 383 Ω 128 This is a public adaptation of a document I wrote for an internal Anthropic audience about a month ago. Thanks to (in alphabetical order) Joshua Batson, Joe Benton, Sam Bowman, Roger Grosse, Jeremy Hadfield, Jared Kaplan, Jan Leike, Jack Lindsey, Monte MacDiarmid, Sam Marks, Fra Mosconi, Chris Olah, Ethan Perez, Sara Price, Ansh Radhakrishnan, Fabien Roger, Buck Shlegeris, Drake Thomas, and Kat
saved by
- Yixiong Hao
- Asher P
- Lydia Nottingham
- Eric Huang
- Emil Ryd
- Jason Hausenloy
- Julian H
- Jo J.
- Arjun Khandelwal
- Neil Rathi
- Will Anderson
- Yannick Mühlhäuser
related reading
- Alignment is not solved but it increasingly looks solvablealigned.substack.com
- AGI Ruin: A List of Lethalities — LessWronglesswrong.com
- Current AIs seem pretty misaligned to me — LessWronglesswrong.com
- Alignment remains a hard, unsolved problem — AI Alignment Forumalignmentforum.org
- Automated Alignment Researchers: Using large language models to scale scalable oversight \ Anthropicanthropic.com
- Can we safely automate alignment research? - Joe Carlsmithjoecarlsmith.com
- Why I’m optimistic about our alignment approachaligned.substack.com
- Shtetl-Optimized >> Blog Archive >> Theory and AI Alignmentscottaaronson.blog
- The Universe from an Intentional Stancecasparoesterheld.com
- PSA: Almost nobody is directly working on superintelligent alignment — LessWronglesswrong.com
- Alignment Is Proven To Be Solvable - by SE Gygesverysane.ai
- A minimal viable product for alignment - by Jan Leikealigned.substack.com