How We Made AI Diplomacy Work
Alex Duffy leads AI training for Every’s consultancy and writes about AI tools and technology in Context Window. Lessons learned taking our game-based AI benchmark from demo to 50,000 live viewers June 25, 2025 When we launched AI Diplomacy earlier this month, we were excited to share what we felt was an innovative AI benchmark, built as a game that anyone could watch and enjoy. The response from readers and the AI research community has been fantastic, and today we’re revealing the details of how Alex Duffy’s labor of love came together.—Michael Reilly Was this newsletter forwarded to you? Sign up to get it in your inbox. "Could you play the game given this context?" That question transformed our barely working AI demo into something 50,000 people watched live on Twitch and millions would see around the world. I built AI Diplomacy along with my friend and developer Tyler Marques because we strongly believe in the power of benchmarks to help us learn about AI and shape its developmen
Skip to content We use analytics and advertising tools by default. You can update this anytime. Privacy Preferences Do Not Sell or Share Privacy Preferences x Manage optional tracking categories. Necessary cookies stay on so the site can function. Analytics Advertising and sharing Cancel Save Preferences How We Made AI Diplomacy Work Midjourney/Every illustration. By Alex Duffy Alex Duffy is the cofounder and CEO of Good Start Labs, and a contributing writer. How We Made AI Diplomacy W ork Lessons learned taking our game-based AI benchmark from demo to 50,000 live viewers Alex Duffy Jun 25, 20
Explore this link on the map →related reading
- Effective context engineering for AI agents \ Anthropicanthropic.com
- 2023 letter | Zhengdongzhengdongwang.com
- CICERO: An AI agent that negotiates, persuades, and cooperates with peopleai.facebook.com
- Building Effective AI Agents \ Anthropicanthropic.com
- Could AI Replace Diplomats? - MI Oasismultipleintelligencesoasis.org
- Building reliable AI agents · parth sareenparthsareen.com
- Import AIjack-clark.net
- [2310.08901] Welfare Diplomacy: Benchmarking Language Model Cooperationarxiv.org
- Context Engineering for AI Agents: Lessons from Building Manusmanus.im
- LLMs Go To Confession, Automated Scientific Research, What Copilot Users Want, and more...deeplearning.ai
- How Long Contexts Faildbreunig.com
- Agent Evaluation: A Detailed Guidecameronrwolfe.substack.com