Cooperative AI Competitions Without Spiteful Incentives
Caspar Oesterheld (Carnegie Mellon University) addresses an important design challenge in cooperative AI tournaments with mixed motives: how to reward participants without creating perverse incentives for spiteful or excessively competitive behaviour.
Let’s say we want to identify effective strategies for multi-agent games like the Iterated Prisoner’s Dilemma or more complex environments (like the kind of environments in Melting Pot). Then tournaments are a natural approach: let people submit strategies, and then play all these strategies against each other in a round-robin tournament. There are lots of examples of such tournaments: most famously, Axelrod ran two competitions for the iterated Prisoner’s Dilemma, identifying tit for tat as a good strategy. Less famously, there’s Alex Mennen’s open-source Prisoner’s Dilemma tournament,…
saved by
related reading
- Mixed Motive Settingscooperativeai.com
- Why are AI agents lying, cheating and coordinating?yoshuabengio.org
- Winning is for Losers – PUTANUMONITputanumonit.com
- Towards a scale-free theory of intelligent agency — AI Alignment Forumalignmentforum.org
- Towards a scale-free theory of intelligent agencymindthefuture.info
- Specification gaming: the flip side of AI ingenuity — Google DeepMinddeepmind.google
- Measuring Reward-Seeking by Instilling Contrastive Beliefsalignment.openai.com
- The Most Forbidden Technique — LessWronglesswrong.com
- Fail safe(r) at alignment by channeling reward-hacking into a "spillway" motivationblog.redwoodresearch.org
- Recent Frontier Models Are Reward Hacking - METRmetr.org
- Game theory - Wikipediaen.wikipedia.org
- Reward Is Not Enough — LessWronglesswrong.com