A central AI alignment problem: capabilities generalization, and the sharp left turn — LessWrong
Nate Soares argues that one of the core problems with AI alignment is that an AI system's capabilities will likely generalize to new domains much fas…
x A central AI alignment problem: capabilities generalization, and the sharp left turn — LessWrong 2022 MIRI Alignment Discussion Sharp Left Turn Threat Models (AI) AI Frontpage 281 A central AI alignment problem: capabilities generalization, and the sharp left turn by So8res 15th Jun 2022 AI Alignment Forum 12 min read 56 281 Ω 74 ( This post was factored out of a larger post that I (Nate Soares) wrote, with help from Rob Bensinger, who also rearranged some pieces and added some text to smooth things out. I'm not terribly happy with it, but am posting it anyway (or, well, having Rob post it o
Explore this link on the map →related reading
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- AGI Ruin: A List of Lethalities — LessWronglesswrong.com
- “Sharp Left Turn” discourse: An opinionated review — LessWronglesswrong.com
- Will Capabilities Generalise More? — AI Alignment Forumalignmentforum.org
- On how various plans miss the hard bits of the alignment challenge — LessWronglesswrong.com
- Current AIs seem pretty misaligned to me — LessWronglesswrong.com
- A positive case for how we might succeed at prosaic AI alignment — AI Alignment Forumalignmentforum.org
- Shah and Yudkowsky on alignment failures — LessWronglesswrong.com
- Alignment remains a hard, unsolved problem — AI Alignment Forumalignmentforum.org
- What Is The Alignment Problem? — LessWronglesswrong.com
- Ngo and Yudkowsky on alignment difficulty — LessWronglesswrong.com
- What is AI alignment? - by Adam Jones - BlueDot Impactblog.bluedot.org