Sections 1 & 2: Introduction, Strategy and Governance — LessWrong
This post is part of the sequence version of the Effective Altruism Foundation's research agenda on Cooperation, Conflict, and Transformative Artificial Intelligence. Transformative artificial intelligence (TAI) may be a key factor in the long-run trajectory of civilization. A growing interdisciplinary community has begun to study how the development of TAI can be made safe and beneficial to sentient life (Bostrom 2014; Russell et al., 2015; OpenAI, 2018; Ortega and Maini, 2018; Dafoe, 2018). We present a research agenda for advancing a critical component of this effort: preventing catastrophic failures of cooperation among TAI systems. By cooperation failures we refer to a broad class of potentially-catastrophic inefficiencies in interactions among TAI-enabled actors. These include destructive conflict; coercion; and social dilemmas (Kollock, 1998; Macy and Flache, 2002) which destroy value over extended periods of time. We introduce cooperation failures at greater length in Section 1
x Sections 1 & 2: Introduction, Strategy and Governance — LessWrong Cooperation, Conflict, and Transformative Artificial Intelligence: A Research Agenda Center on Long-Term Risk (CLR) Risks of Astronomical Suffering (S-risks) Coordination / Cooperation Game Theory Research Agendas AI Frontpage 35 Sections 1 & 2: Introduction, Strategy and Governance by JesseClifton 17th Dec 2019 AI Alignment Forum 16 min read 8 35 Ω 9 This post is part of the sequence version of the Effective Altruism Foundation's research agenda on Cooperation, Conflict, and Transformative Artificial Intelligence . 1 Introduc
Explore this link on the map →related reading
- Gradual Paths to Collective Flourishing — LessWronglesswrong.com
- What failure looks like — LessWronglesswrong.com
- Where I agree and disagree with Eliezer — LessWronglesswrong.com
- Making deals with early schemers — LessWronglesswrong.com
- Solipsistic Superintelligence is Unlikely to be Cooperativearxiv.org
- Another (outer) alignment failure story — AI Alignment Forumalignmentforum.org
- Towards a scale-free theory of intelligent agency — AI Alignment Forumalignmentforum.org
- [2602.12316] GT-HarmBench: Benchmarking AI Safety Risks Through the Lens of Game Theoryarxiv.org
- Concrete Projects in AGI Preparednessforethought.org
- Power Lies Trembling: a three-book review — LessWronglesswrong.com
- AI Could Defeat All Of Us Combined — LessWronglesswrong.com
- AI Pause Will Likely Backfire — EA Forumforum.effectivealtruism.org