Safe Pareto Improvements Research Agenda — Center on Long-Term Risk
Safe Pareto improvements (SPIs) are modifications to agents’ bargaining strategies that make all parties better off, regardless of their original strategies.
Executive summary Safe Pareto improvements (SPIs) are ways of changing agents' bargaining strategies that make all parties better off, regardless of their original strategies. SPIs are an unusually robust approach to preventing catastrophic conflict between AI systems, especially AIs capable of credible commitments. This is because SPIs can reduce the costs of conflict without shifting bargaining power, or requiring agents to agree on what counts as "fair". Despite their appeal, SPIs aren't guaranteed to be adopted. AIs or humans in the loop might lock in SPI-incompatible commitments, or…
saved by
related reading
- A gap in the theoretical justification for surrogate goals and safe Pareto improvements – The Universe from an Intentional Stancecasparoesterheld.com
- Responsible Scaling Policy v3 — LessWronglesswrong.com
- Cooperation, Conflict, and Transformative Artificial Intelligence: A Research Agenda — Center on Long-Term Risklongtermrisk.org
- AI Control: Improving Safety Despite Intentional Subversion — LessWronglesswrong.com
- Towards a scale-free theory of intelligent agency — AI Alignment Forumalignmentforum.org
- FLI AI Safety Research Landscape - Extended - v0.43futureoflife.org
- Solipsistic Superintelligence is Unlikely to be Cooperativearxiv.org
- Coasean Bargaining at Scaleblog.cosmos-institute.org
- Spring 2026 Projects - SPARsparai.org
- Patterns and problems in multiagent systemsanthropic.com
- Making deals with early schemers — LessWronglesswrong.com
- Optimality is the tiger, and agents are its teeth — LessWronglesswrong.com