1. The CAST Strategy — LessWrong
(TLDR for this section, since it’s 101 stuff that many readers will have already grokked: Misuse vs Mistake; Principal-Agent problem; Omohundro Drives; we need deep safety measures in addition to mundane methods. Jump to “Sleepy-Bot” if all that seems familiar.) Earth is in peril. Humanity is on the verge of building machines capable of intelligent action that outstrips our collective wisdom. These superintelligent artificial general intelligences (“AGIs”) are almost certain to radically transform the world, perhaps very quickly, and likely in ways that we consider catastrophic, such as driving humanity to extinction. During this pivotal period, our peril manifests in two forms. The most obvious peril is that of misuse. An AGI which is built to serve the interests of one person or party, such as jihadists or tyrants, may harm humanity as a whole (e.g. by producing bioweapons or mind-control technology). Power tends to corrupt, and if a small number of people have power over armies of m
x 1. The CAST Strategy — LessWrong CAST: Corrigibility As Singular Target Corrigibility AI Frontpage 58 1. The CAST Strategy by Max Harms 7th Jun 2024 AI Alignment Forum 46 min read 27 58 Ω 23 (Part 1 of the CAST sequence ) AI Risk Introduction (TLDR for this section, since it’s 101 stuff that many readers will have already grokked: Misuse vs Mistake; Principal-Agent problem; Omohundro Drives; we need deep safety measures in addition to mundane methods. Jump to “Sleepy-Bot” if all that seems familiar.) Earth is in peril. Humanity is on the verge of building machines capable of intelligent acti
Explore this link on the map →related reading
- Let's See You Write That Corrigibility Tag — LessWronglesswrong.com
- Dario Amodei — The Adolescence of Technologydarioamodei.com
- Terrified Comments on Corrigibility in Claude's Constitution — LessWronglesswrong.com
- 1a3orn's Shortform — LessWronglesswrong.com
- AGI Ruin: A List of Lethalities — LessWronglesswrong.com
- ROGUE:arxiv.org
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- Commentary on AGI Safety from First Principles — AI Alignment Forumalignmentforum.org
- The hot mess theory of AI misalignment: More intelligent agents behave less coherently | Jascha’s blogsohl-dickstein.github.io
- [2502.15657] Superintelligent Agents Pose Catastrophic Risks: Can Scientist AI Offer a Safer Path?arxiv.org
- AI as a science, and three obstacles to alignment strategies — LessWronglesswrong.com
- A positive case for how we might succeed at prosaic AI alignment — AI Alignment Forumalignmentforum.org