Reevaluating Policy Gradient Methods for Imperfect-Information Games
This is experimental HTML to improve accessibility. We invite you to report rendering errors. Use Alt+Y to toggle on accessible reporting links and Alt+Shift+Y to toggle off. Learn more about this project and help improve conversions. In the past decade, motivated by the putative failure of naive self-play deep reinforcement learning (DRL) in adversarial imperfect-information games, researchers have developed numerous DRL algorithms based on fictitious play (FP), double oracle (DO), and counterfactual regret minimization (CFR). In light of recent results of the magnetic mirror descent algorithm, we hypothesize that simpler generic policy gradient methods like PPO are competitive with or superior to these FP-, DO-, and CFR-based DRL approaches. To facilitate the resolution of this hypothesis, we implement and release the first broadly accessible exact exploitability computations for four large games. Using these games, we conduct the largest-ever exploitability comparison of DRL algori
Reevaluating Policy Gradient Methods for Imperfect-Information Games Max Rudolph University of Texas at Austin Nathan Lichtlé University of California, Berkeley Sobhan Mohammadpour Massachusetts Institute of Technology Alexandre Bayen University of California, Berkeley J. Zico Kolter Carnegie Mellon University Amy Zhang University of Texas at Austin Gabriele Farina Massachusetts Institute of Technology Eugene Vinitsky NYU Tandon School of Engineering Samuel Sokota Carnegie Mellon University Abstract In the past decade, motivated by the putative failure of naive self-play deep reinforcement lea
Explore this link on the map →related reading
- Policy Gradient Algorithms | Lil'Loglilianweng.github.io
- A (Long) Peek into Reinforcement Learning | Lil'Loglilianweng.github.io
- The 37 Implementation Details of Proximal Policy Optimization · The ICLR Blog Trackiclr-blog-track.github.io
- Key Papers in Deep RL - Spinning Up documentationspinningup.openai.com
- Reinforcement Learning in Newcomblike Problemsproceedings.neurips.cc
- State of RL for reasoning LLMs | A. Weersaweers.de
- Deep Reinforcement Learning: Pong from Pixelskarpathy.github.io
- Learning Beyond Gradientstrinkle23897.github.io
- [1709.06560] Deep Reinforcement Learning that Mattersarxiv.org
- Why do Policy Gradient Methods work so well in Cooperative MARL? Evidence from Policy Representation – The Berkeley Artificial Intelligence Research Blogbair.berkeley.edu
- Natural Deception with RL - Rajan Agarwalrajan.sh
- [1707.06347] Proximal Policy Optimization Algorithmsarxiv.org