[2310.08901] Welfare Diplomacy: Benchmarking Language Model Cooperation
Abstract:The growing capabilities and increasingly widespread deployment of AI systems necessitate robust benchmarks for measuring their cooperative capabilities. Unfortunately, most multi-agent benchmarks are either zero-sum or purely cooperative, providing limited opportunities for such measurements. We introduce a general-sum variant of the zero-sum board game Diplomacy -- called Welfare Diplomacy -- in which players must balance investing in military conquest and domestic welfare. We argue that Welfare Diplomacy facilitates both a clearer assessment of and stronger training incentives for cooperative capabilities. Our contributions are: (1) proposing the Welfare Diplomacy rules and implementing them via an open-source Diplomacy engine; (2) constructing baseline agents using zero-shot prompted language models; and (3) conducting experiments where we find that baselines using state-of-the-art models attain high social welfare but are exploitable. Our work aims to promote societal safety by aiding researchers in developing and assessing multi-agent AI systems. Code to evaluate Welfare Diplomacy and reproduce our experiments is available at this https URL.
[2310.08901] Welfare Diplomacy: Benchmarking Language Model Cooperation Skip to main content arXiv is now an independent nonprofit! Learn more × Search arXiv Press Enter to search · Advanced search --> Computer Science > Multiagent Systems arXiv:2310.08901 (cs) [Submitted on 13 Oct 2023] Title: Welfare Diplomacy: Benchmarking Language Model Cooperation Authors: Gabriel Mukobi , Hannah Erlebach , Niklas Lauffer , Lewis Hammond , Alan Chan , Jesse Clifton View a PDF of the paper titled Welfare Diplomacy: Benchmarking Language Model Cooperation, by Gabriel Mukobi and 5 other authors
Explore this link on the map →related reading
- [2602.12316] GT-HarmBench: Benchmarking AI Safety Risks Through the Lens of Game Theoryarxiv.org
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- Solipsistic Superintelligence is Unlikely to be Cooperativearxiv.org
- [2411.00114] Project Sid: Many-agent simulations toward AI civilizationarxiv.org
- Agent Evaluation: A Detailed Guidecameronrwolfe.substack.com
- PostTrainBenchposttrainbench.com
- [2303.13360] Towards the Scalable Evaluation of Cooperativeness in Language Modelsarxiv.org
- CICERO: An AI agent that negotiates, persuades, and cooperates with peopleai.facebook.com
- Natural Deception with RL - Rajan Agarwalrajan.sh
- Spring 2026 Projects - SPARsparai.org
- Towards a scale-free theory of intelligent agency — AI Alignment Forumalignmentforum.org
- Escalation Risks from Language Models in Military and Diplomatic Decision-Makingarxiv.org