VGC-Bench: Towards Mastering Diverse Team Strategies in Competitive Pokémon
Abstract:Developing AI agents that can robustly adapt to varying strategic landscapes without retraining is a central challenge in multi-agent learning. Pokémon Video Game Championships (VGC) is a domain with a vast space of approximately $10^{139}$ team configurations, far larger than those of other games such as Chess, Go, Poker, StarCraft, or Dota. The combinatorial nature of team building in Pokémon VGC causes optimal strategies to vary substantially depending on both the controlled team and the opponent's team, making generalization uniquely challenging. To advance research on this problem, we introduce VGC-Bench: a benchmark that provides critical infrastructure, standardizes evaluation protocols, and supplies a human-play dataset of over 700,000 battle logs and a range of baseline agents based on heuristics, large language models, behavior cloning, and multi-agent reinforcement learning with empirical game-theoretic methods such as self-play, fictitious play, and double oracle. In the restricted setting where an agent is trained and evaluated in a mirror match with a single team configuration, our methods can win against a professional VGC competitor. We repeat this training and evaluation with progressively larger team sets and find that as the number of teams increases, the best-performing algorithm in the single-team setting has worse performance and is more exploitable, but has improved generalization to unseen teams. Our code and dataset are open-sourced at this https URL and this https URL.
# link_20t20k9yyvh.pdf ## Metadata - PDFFormatVersion=1.7 - Language=en - IsLinearized=false - IsAcroFormPresent=false - IsXFAPresent=false - IsCollectionPresent=false - IsSignaturesPresent=false - Author=Cameron Angliss; Jiaxun Cui; Jiaheng Hu; Arrasy Rahman; Peter Stone - CreationDate=D:20260114015017+00'00' - Creator=arXiv GenPDF (tex2pdf:57610bf) - Custom.DOI=https://doi.org/10.48550/arXiv.2506.10326 - Custom.License=http://creativecommons.org/licenses/by/4.0/ - Custom.PTEX.Fullbanner=This is pdfTeX, Version 3.141592653-2.6-1.40.28 (TeX Live 2025) kpathsea version 6.4.1 - Custom.arXivID=ht
Explore this link on the map →saved by
related reading
- The 37 Implementation Details of Proximal Policy Optimization · The ICLR Blog Trackiclr-blog-track.github.io
- Composer2.pdfcursor.com
- AI Models for Pokemon Games - Kevin Lukevinlu.ai
- How DeepMind's Generally Capable Agents Were Trained — LessWronglesswrong.com
- HANABI – np – ( ´ ▽ ` )ノnphard.io
- [AN #159]: Building agents that know how to experiment, by training on procedurally generated games — AI Alignment Forumalignmentforum.org
- Humans Still Beat AI in the Long Horizon: Revisiting Test-Time Scaling in the Agent Era | Qiuyang Mangjoyemang33.github.io
- Generally capable agents emerge from open-ended play — Google DeepMinddeepmind.google
- Learning Beyond Gradientstrinkle23897.github.io
- [2605.22748] Superhuman Safe and Agile Racing through Multi-Agent Reinforcement Learningarxiv.org
- DeepMind: Generally capable agents emerge from open-ended play — LessWronglesswrong.com
- Seth Karten on X: "https://t.co/mTE4Dta508" / Xx.com