Box inversion revisited — AI Alignment Forum
Box inversion hypothesis is a proposed correspondence between problems with AI systems studied in approaches like agent foundations, and problems wit…
x Box inversion revisited — AI Alignment Forum Agent Foundations AI Services (CAIS) AI Frontpage 19 Box inversion revisited by Jan_Kulveit 7th Nov 2023 9 min read 3 19 Box inversion hypothesis is a proposed correspondence between problems with AI systems studied in approaches like agent foundations , and problems with AI ecosystems, studied in various views on AI safety expecting multipolar, complex worlds, like CAIS. This is an updated and improved introduction to the idea. Cartoon explanation In the classic -"superintelligence in a box" - picture, we worry about an increasingly powerful AGI,
Explore this link on the map →related reading
- Dario Amodei — The Adolescence of Technologydarioamodei.com
- AGI Ruin: A List of Lethalities — LessWronglesswrong.com
- The Possessed Machines: Dostoevsky's Demons and the Coming AGI Catastrophepossessedmachines.com
- Core views on AI safety: When, why, what, and how \ Anthropicanthropic.com
- Current AIs seem pretty misaligned to me — LessWronglesswrong.com
- Core views on AI safety: When, why, what, and how \ Anthropicanthropic.com
- What failure looks like — LessWronglesswrong.com
- A positive case for how we might succeed at prosaic AI alignment — AI Alignment Forumalignmentforum.org
- AI safety - Wikipediaen.wikipedia.org
- Commentary on AGI Safety from First Principles — AI Alignment Forumalignmentforum.org
- AGI safety from first principles: Introduction — LessWronglesswrong.com
- A Field Guide to AI Safety—Asteriskasteriskmag.com