Solipsistic Superintelligence is Unlikely to be Cooperative
Abstract:AI's central challenge is shifting from capability to coexistence. The dominant paradigm in AI research focuses on developing powerful agents that treat the world as an exogenous and stationary source of feedback. We contend that superintelligence, an extremely capable task solver, born out of such a solipsistic approach to AI design, is unlikely to be cooperative. Deploying AI systems induces endogenous non-stationarity, resulting in a train-test-deploy gap where historical distributions diverge from the deployment context. We refer to this as the self-undermining property of unilateral optimization. Closing this gap requires AI that participates in cooperation: the equilibrium-selection process through which multiple actors navigate their interdependence. We call for a non-solipsistic research paradigm that treats this interdependence as a core design principle rather than approaching cooperation as a task to solve. This entails building dynamic evaluation testbeds involving adaptive counterparties, treating institutions as design primitives, and preserving human agency as a structural feature of the systems we build.
# link_3ayt9a49b6.pdf ## Metadata - PDFFormatVersion=1.7 - IsLinearized=false - IsAcroFormPresent=false - IsXFAPresent=false - IsCollectionPresent=false - IsSignaturesPresent=false - Author=Rakshit S Trivedi; Natasha Jaques; Logan Cross; Alexander Sasha Vezhnevets; Joel Z Leibo - Creator=arXiv GenPDF (tex2pdf:a6404ea) - Custom.DOI=https://doi.org/10.48550/arXiv.2606.03237 - Custom.License=http://creativecommons.org/licenses/by/4.0/ - Custom.PTEX.Fullbanner=This is pdfTeX, Version 3.141592653-2.6-1.40.28 (TeX Live 2025) kpathsea version 6.4.1 - Custom.arXivID=https://arxiv.org/abs/2606.03237v1
Explore this link on the map →saved by
related reading
- My AI Opinions - by Scott Alexander - Astral Codex Tensubstack.com
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- The hot mess theory of AI misalignment: More intelligent agents behave less coherently | Jascha’s blogsohl-dickstein.github.io
- [2502.15657] Superintelligent Agents Pose Catastrophic Risks: Can Scientist AI Offer a Safer Path?arxiv.org
- A positive case for how we might succeed at prosaic AI alignment — AI Alignment Forumalignmentforum.org
- ⿻ Symbiogenesis vs. Convergent Consequentialism — LessWronglesswrong.com
- Book Review: Reframing Superintelligence | Slate Star Codexslatestarcodex.com
- Concrete Projects in AGI Preparednessforethought.org
- Superintelligence: The Idea That Eats Smart Peopleidlewords.com
- IIIc. Superalignment - SITUATIONAL AWARENESSsituational-awareness.ai
- Towards a scale-free theory of intelligent agency — AI Alignment Forumalignmentforum.org
- A Field Guide to AI Safety—Asteriskasteriskmag.com