Treacherous turns in the wild
…one idea for how to ensure superintelligence safety… is that we validate the safety of a superintelligent AI empirically by observing its behavior while it is in a controlled, limited environment (a “sandbox”) and that we only let the AI out of the box if we see it behaving in a friendly, cooperative, responsible manner. The flaw in this idea is that behaving nicely while in the box is a convergent instrumental goal for friendly and unfriendly AIs alike. An unfriendly AI of sufficient intelligence realizes that its unfriendly final goals will be best realized if it behaves in a friendly manner initially, so that it will be let out of the box. It will only start behaving in a way that reveals its unfriendly nature when it no longer matters whether we find out; that is, when the AI is strong enough that human opposition is ineffectual. Some people have told me they think this is unrealistic, apparently even for a machine superintelligence far more capable than any current AI system. But
…one idea for how to ensure superintelligence safety… is that we validate the safety of a superintelligent AI empirically by observing its behavior while it is in a controlled, limited environment (a “sandbox”) and that we only let the AI out of the box if we see it behaving in a friendly, cooperative, responsible manner. The flaw in this idea is that behaving nicely while in the box is a convergent instrumental goal for friendly and unfriendly AIs alike. An unfriendly AI of sufficient intelligence realizes that its unfriendly final goals will be best realized if it behaves in a friendly manne
Explore this link on the map →saved by
related reading
- Dario Amodei — The Adolescence of Technologydarioamodei.com
- The case for ensuring that powerful AIs are controlled — LessWronglesswrong.com
- Unfalsifiable stories of doom | Mechanize, Inc.mechanize.work
- AI Control: Improving Safety Despite Intentional Subversion — LessWronglesswrong.com
- Which Future?michaelnotebook.com
- Superintelligence FAQ — LessWronglesswrong.com
- Superintelligence: The Idea That Eats Smart Peopleidlewords.com
- The case for ensuring that powerful AIs are controlledblog.redwoodresearch.org
- A Field Guide to AI Safety—Asteriskasteriskmag.com
- AI Control: Improving Safety Despite Intentional Subversion — AI Alignment Forumalignmentforum.org
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- [2502.15657] Superintelligent Agents Pose Catastrophic Risks: Can Scientist AI Offer a Safer Path?arxiv.org