Alignment remains a hard, unsolved problem — LessWrong
This is a public adaptation of a document I wrote for an internal Anthropic audience about a month ago. Thanks to (in alphabetical order) Joshua Bats…
x Alignment remains a hard, unsolved problem — LessWrong AI Curated 2025 Top Fifty: 25 % 383 Alignment remains a hard, unsolved problem by evhub 27th Nov 2025 AI Alignment Forum 16 min read 98 383 Ω 128 This is a public adaptation of a document I wrote for an internal Anthropic audience about a month ago. Thanks to (in alphabetical order) Joshua Batson, Joe Benton, Sam Bowman, Roger Grosse, Jeremy Hadfield, Jared Kaplan, Jan Leike, Jack Lindsey, Monte MacDiarmid, Sam Marks, Fra Mosconi, Chris Olah, Ethan Perez, Sara Price, Ansh Radhakrishnan, Fabien Roger, Buck Shlegeris, Drake Thomas, and Kat
Explore this link on the map →saved by
- Yixiong Hao
- Asher P
- Lydia Nottingham
- Eric Huang
- Emil Ryd
- Jason Hausenloy
- Julian H
- Jo J.
- Arjun Khandelwal
- Neil Rathi
- Will Anderson
- 6194
related reading
- AGI Ruin: A List of Lethalities — LessWronglesswrong.com
- Current AIs seem pretty misaligned to me — LessWronglesswrong.com
- Alignment remains a hard, unsolved problem — AI Alignment Forumalignmentforum.org
- Automated Alignment Researchers: Using large language models to scale scalable oversight \ Anthropicanthropic.com
- Why I’m optimistic about our alignment approachaligned.substack.com
- PSA: Almost nobody is directly working on superintelligent alignment — LessWronglesswrong.com
- A minimal viable product for alignment - by Jan Leikealigned.substack.com
- Teaching Claude why \ Anthropicanthropic.com
- A positive case for how we might succeed at prosaic AI alignment — AI Alignment Forumalignmentforum.org
- Why I’m optimistic about our alignment approachaligned.substack.com
- Ngo and Yudkowsky on alignment difficulty — LessWronglesswrong.com
- Can we safely automate alignment research? - Joe Carlsmithjoecarlsmith.com