[1903.08894] Towards Characterizing Divergence in Deep Q-Learning
arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs. arXiv Operational Status Get status notifications via email or slack
[1903.08894] Towards Characterizing Divergence in Deep Q-Learning Skip to main content arXiv is now an independent nonprofit! Learn more × Search arXiv Press Enter to search · Advanced search --> Computer Science > Machine Learning arXiv:1903.08894 (cs) [Submitted on 21 Mar 2019] Title: Towards Characterizing Divergence in Deep Q-Learning Authors: Joshua Achiam , Ethan Knight , Pieter Abbeel View a PDF of the paper titled Towards Characterizing Divergence in Deep Q-Learning, by Joshua Achiam and 2 other authors View PDF Abstract: Deep Q-Learning (DQL), a family of temporal differe
Explore this link on the map →related reading
- A (Long) Peek into Reinforcement Learning | Lil'Loglilianweng.github.io
- Deep Q-Networks Explained — LessWronglesswrong.com
- Key Papers in Deep RL - Spinning Up documentationspinningup.openai.com
- Q-learning is not yet scalableseohong.me
- Technical Note: Q-Learning | Machine Learning | Springer Nature Linklink.springer.com
- Policy Gradient Algorithms | Lil'Loglilianweng.github.io
- [1709.06560] Deep Reinforcement Learning that Mattersarxiv.org
- Q-learning - Wikipediaen.wikipedia.org
- The Inverted Pendulum Problem with Deep Reinforcement Learning | by Saif Uddin Mahmud | Dabbler in Destress | Mediummedium.com
- Seohong Park on X: "We scaled up an "alternative" paradigm in RL: *divide and conquer*. Compared to Q-learning (TD learning), divide and conquer can naturally scale to much longer horizons. Blog post: https://t.co/xtXBzya0bI Paper: https://t.co/nqYkLucsWu ↓ https://t.co/XCdgUzaLxF" / Xx.com
- Part 2: Kinds of RL Algorithms - Spinning Up documentationspinningup.openai.com
- RL's Deadly Triad Meets Optimization | helen quhelenqu.com