What Is The Alignment Problem? — LessWrong
lesswrong.com · 17,467 words · saved by 1 readers
So we want to align future AGIs. Ultimately we’d like to align them to human values, but in the shorter term we might start with other targets, like…
x What Is The Alignment Problem? — LessWrong AI Curated 2025 Top Fifty: 16 % 184 What Is The Alignment Problem? by johnswentworth 16th Jan 2025 AI Alignment Forum 30 min read 49 184 Ω 65 So we want to align future AGIs. Ultimately we’d like to align them to human values, but in the shorter term we might start with other targets, like e.g. corrigibility . That problem description all makes sense on a hand-wavy intuitive level, but once we get concrete and dig into technical details… wait, what exactly is the goal again? When we say we want to “align AGI”, what does that mean? And what about the
related reading
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- AGI Ruin: A List of Lethalities — LessWronglesswrong.com
- ALIGNMENT - by vincent huang - a slice of my mindmindslice.substack.com
- What Does It Mean to Align AI With Human Values? | Quanta Magazinequantamagazine.org
- The Artificiality of Alignmentjoinreboot.org
- Ngo and Yudkowsky on alignment difficulty — LessWronglesswrong.com
- What is AI alignment? - by Adam Jones - BlueDot Impactblog.bluedot.org
- Shah and Yudkowsky on alignment failures — LessWronglesswrong.com
- On how various plans miss the hard bits of the alignment challenge — LessWronglesswrong.com
- Alignment Is Proven To Be Solvable - by SE Gygesverysane.ai
- A positive case for how we might succeed at prosaic AI alignment — AI Alignment Forumalignmentforum.org
- Explore: AI alignmentarbital.greaterwrong.com