flâneur — a map of the web's best reading

Reevaluating Policy Gradient Methods for Imperfect-Information Games

arxiv.org · 14,699 words · saved by 1 readers

This is experimental HTML to improve accessibility. We invite you to report rendering errors. Use Alt+Y to toggle on accessible reporting links and Alt+Shift+Y to toggle off. Learn more about this project and help improve conversions. In the past decade, motivated by the putative failure of naive self-play deep reinforcement learning (DRL) in adversarial imperfect-information games, researchers have developed numerous DRL algorithms based on fictitious play (FP), double oracle (DO), and counterfactual regret minimization (CFR). In light of recent results of the magnetic mirror descent algorithm, we hypothesize that simpler generic policy gradient methods like PPO are competitive with or superior to these FP-, DO-, and CFR-based DRL approaches. To facilitate the resolution of this hypothesis, we implement and release the first broadly accessible exact exploitability computations for four large games. Using these games, we conduct the largest-ever exploitability comparison of DRL algori

Reevaluating Policy Gradient Methods for Imperfect-Information Games Max Rudolph University of Texas at Austin Nathan Lichtlé University of California, Berkeley Sobhan Mohammadpour Massachusetts Institute of Technology Alexandre Bayen University of California, Berkeley J. Zico Kolter Carnegie Mellon University Amy Zhang University of Texas at Austin Gabriele Farina Massachusetts Institute of Technology Eugene Vinitsky NYU Tandon School of Engineering Samuel Sokota Carnegie Mellon University Abstract In the past decade, motivated by the putative failure of naive self-play deep reinforcement lea

Explore this link on the map →

related reading