How might we align transformative AI if it’s developed very soon? — LessWrong
This post gives my understanding of what the set of available strategies for aligning transformative AI would be if it were developed very soon, and why they might or might not work. It is heavily based on conversations with Paul Christiano, Ajeya Cotra and Carl Shulman, and its background assumptions correspond to the arguments Ajeya makes in this piece (abbreviated as “Takeover Analysis”). I premise this piece on a nearcast in which a major AI company (“Magma,” following Ajeya’s terminology) has good reason to think that it can develop transformative AI very soon (within a year), using what Ajeya calls “human feedback on diverse tasks” (HFDT) - and has some time (more than 6 months, but less than 2 years) to set up special measures to reduce the risks of misaligned AI before there’s much chance of someone else deploying transformative AI. I will discuss: A few of the uncertain factors that seem most important to me are (a) how cautious key actors are; (b) whether AI systems have big,
x How might we align transformative AI if it’s developed very soon? — LessWrong AI Timelines AI Frontpage 145 How might we align transformative AI if it’s developed very soon? by HoldenKarnofsky 29th Aug 2022 AI Alignment Forum 54 min read 55 145 Ω 52 This post is part of my AI strategy nearcasting series : trying to answer key strategic questions about transformative AI, under the assumption that key events will happen very soon, and/or in a world that is otherwise very similar to today's. This post gives my understanding of what the set of available strategies for aligning transformative AI
Explore this link on the map →related reading
- Nearcast-based "deployment problem" analysis — AI Alignment Forumalignmentforum.org
- Dario Amodei — The Adolescence of Technologydarioamodei.com
- Without specific countermeasures, the easiest path to transformative AI likely leads to AI takeover — AI Alignment Forumalignmentforum.org
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- Current AIs seem pretty misaligned to me — LessWronglesswrong.com
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- How we could stumble into AI catastrophecold-takes.com
- Without specific countermeasures, the easiest path to transformative AI likely leads to AI takeover — LessWronglesswrong.com
- The case for ensuring that powerful AIs are controlled — LessWronglesswrong.com
- Off Target | CNAScnas.org
- A positive case for how we might succeed at prosaic AI alignment — AI Alignment Forumalignmentforum.org
- Another (outer) alignment failure story — AI Alignment Forumalignmentforum.org