Another (outer) alignment failure story - AI Alignment Forum
META This is a story where the alignment problem is somewhat harder than I expect, society handles AI more competently than I expect, and the outcome is worse than I expect. It also involves inner al…
x Another (outer) alignment failure story — AI Alignment Forum Best of LessWrong 2021 Threat Models (AI) Outer Alignment AI Risk AI Curated 76 Another (outer) alignment failure story by paulfchristiano 7th Apr 2021 14 min read 39 76 Meta This is a story where the alignment problem is somewhat harder than I expect, society handles AI more competently than I expect, and the outcome is worse than I expect. It also involves inner alignment turning out to be a surprisingly small problem. Maybe the story is 10-20th percentile on each of those axes. At the end I’m going to go through some salient way
Explore this link on the map →saved by
related reading
- What failure looks like — LessWronglesswrong.com
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- Dario Amodei — The Adolescence of Technologydarioamodei.com
- Current AIs seem pretty misaligned to me — LessWronglesswrong.com
- What failure looks like — AI Alignment Forumalignmentforum.org
- AGI Ruin: A List of Lethalities — LessWronglesswrong.com
- Shah and Yudkowsky on alignment failures — LessWronglesswrong.com
- Where I agree and disagree with Eliezer — LessWronglesswrong.com
- Current AIs seem pretty misaligned to meblog.redwoodresearch.org
- What Failure Looks Like: Distilling the Discussion — LessWronglesswrong.com
- A positive case for how we might succeed at prosaic AI alignment — AI Alignment Forumalignmentforum.org
- We're already in AI takeoff — LessWronglesswrong.com