A positive case for how we might succeed at prosaic AI alignment — AI Alignment Forum
This post is my attempt at something like a response to Eliezer Yudkowsky’s recent discussion on AGI interventions. I tend to be relatively pessimistic overall about humanity’s chances at avoiding AI existential risk. Contrary to some others that share my pessimism, however—Eliezer Yudkowsky, in particular—I believe that there is a clear path forward for how we might succeed within the current prosaic paradigm (that is, the current machine learning paradigm) that looks plausible and has no fundamental obstacles. In the comments on Eliezer’s discussion, this point about whether there exists any coherent story for prosaic AI alignment success came up multiple times. From Rob Bensinger: I think it's pretty important here to focus on the object-level. Even if you think the goodness of these particular research directions isn't cruxy (because there's a huge list of other things you find promising, and your view is mainly about the list as a whole rather than about any particular items on it
x A positive case for how we might succeed at prosaic AI alignment — AI Alignment Forum AI Success Models Outer Alignment Prosaic Alignment AI Frontpage 43 A positive case for how we might succeed at prosaic AI alignment by evhub 16th Nov 2021 8 min read 46 43 This post is my attempt at something like a response to Eliezer Yudkowsky’s recent discussion on AGI interventions . I tend to be relatively pessimistic overall about humanity’s chances at avoiding AI existential risk. Contrary to some others that share my pessimism, however— Eliezer Yudkowsky, in particular —I believe that there is a cl
Explore this link on the map →saved by
related reading
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- AGI Ruin: A List of Lethalities — LessWronglesswrong.com
- Unfalsifiable stories of doom | Mechanize, Inc.mechanize.work
- Current AIs seem pretty misaligned to me — LessWronglesswrong.com
- Where I agree and disagree with Eliezer — LessWronglesswrong.com
- Shah and Yudkowsky on alignment failures — LessWronglesswrong.com
- Commentary on AGI Safety from First Principles — AI Alignment Forumalignmentforum.org
- Critical review of Christiano's disagreements with Yudkowsky — LessWronglesswrong.com
- Current AIs seem pretty misaligned to meblog.redwoodresearch.org
- Why AI alignment could be hard with modern deep learningcold-takes.com
- Optimality is the tiger, and agents are its teeth — LessWronglesswrong.com
- Alignment remains a hard, unsolved problem — AI Alignment Forumalignmentforum.org