[1709.06560] Deep Reinforcement Learning that Matters
Abstract:In recent years, significant progress has been made in solving challenging problems across various domains using deep reinforcement learning (RL). Reproducing existing work and accurately judging the improvements offered by novel methods is vital to sustaining this progress. Unfortunately, reproducing results for state-of-the-art deep RL methods is seldom straightforward. In particular, non-determinism in standard benchmark environments, combined with variance intrinsic to the methods, can make reported results tough to interpret. Without significance metrics and tighter standardization of experimental reporting, it is difficult to determine whether improvements over the prior state-of-the-art are meaningful. In this paper, we investigate challenges posed by reproducibility, proper experimental techniques, and reporting procedures. We illustrate the variability in reported metrics and results when comparing against common baselines and suggest guidelines to make future results in deep RL more reproducible. We aim to spur discussion about how to ensure continued progress in the field by minimizing wasted effort stemming from results that are non-reproducible and easily misinterpreted.
# link_1v8d61q7pe4.pdf ## Metadata - PDFFormatVersion=1.5 - IsLinearized=false - IsAcroFormPresent=false - IsXFAPresent=false - IsCollectionPresent=false - IsSignaturesPresent=false - Author=Peter Henderson, Riashat Islam, Philip Bachman, Joelle Pineau, Doina Precup, David Meger - CreationDate=D:20190131020749Z - Creator=TeX - ModDate=D:20190131020749Z - Custom.PTEX.Fullbanner=This is pdfTeX, Version 3.14159265-2.6-1.40.17 (TeX Live 2016) kpathsea version 6.2.2 - Producer=pdfTeX-1.40.17 - Title=Deep Reinforcement Learning that Matters - Trapped=False ## Contents ### Page 1 Deep Reinforcemen
Explore this link on the map →saved by
related reading
- The 37 Implementation Details of Proximal Policy Optimization · The ICLR Blog Trackiclr-blog-track.github.io
- Debugging Reinforcement Learning Systemsandyljones.com
- Key Papers in Deep RL - Spinning Up documentationspinningup.openai.com
- Deep Reinforcement Learning Doesn't Work Yetalexirpan.com
- A (Long) Peek into Reinforcement Learning | Lil'Loglilianweng.github.io
- Part 1: Key Concepts in RL - Spinning Up documentationspinningup.openai.com
- State of RL for reasoning LLMs | A. Weersaweers.de
- [1707.06347] Proximal Policy Optimization Algorithmsarxiv.org
- Q-learning is not yet scalableseohong.me
- The Decade of Deep Learning | Leo Gaobmk.sh
- Deep Q-Networks Explained — LessWronglesswrong.com
- Deep Reinforcement Learning: Pong from Pixelskarpathy.github.io