✳flâneur — a map of the web's best reading
What Is The Alignment Problem? — LessWrong
lesswrong.com · 17,467 words · saved by 1 readers
So we want to align future AGIs. Ultimately we’d like to align them to human values, but in the shorter term we might start with other targets, like…
x What Is The Alignment Problem? — LessWrong AI Curated 2025 Top Fifty: 16 % 184 What Is The Alignment Problem? by johnswentworth 16th Jan 2025 AI Alignment Forum 30 min read 49 184 Ω 65 So we want to align future AGIs. Ultimately we’d like to align them to human values, but in the shorter term we might start with other targets, like e.g. corrigibility . That problem description all makes sense on a hand-wavy intuitive level, but once we get concrete and dig into technical details… wait, what exactly is the goal again? When we say we want to “align AGI”, what does that mean? And what about the
Explore this link on the map →related reading
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- AGI Ruin: A List of Lethalities — LessWronglesswrong.com
- ALIGNMENT - by vincent huang - a slice of my mindmindslice.substack.com
- What Does It Mean to Align AI With Human Values? | Quanta Magazinequantamagazine.org
- Ngo and Yudkowsky on alignment difficulty — LessWronglesswrong.com
- Shah and Yudkowsky on alignment failures — LessWronglesswrong.com
- On how various plans miss the hard bits of the alignment challenge — LessWronglesswrong.com
- The hot mess theory of AI misalignment: More intelligent agents behave less coherently | Jascha’s blogsohl-dickstein.github.io
- A positive case for how we might succeed at prosaic AI alignment — AI Alignment Forumalignmentforum.org
- Why Agent Foundations? An Overly Abstract Explanation — LessWronglesswrong.com
- The self-unalignment problem — LessWronglesswrong.com
- AI as a science, and three obstacles to alignment strategies — LessWronglesswrong.com