Exploration via Elliptical Episodic Bonuses | HTML5
In recent years, a number of reinforcement learning (RL) methods have been proposed to explore complex environments which differ across episodes. In this work, we show that the effectiveness of these methods critically relies on a count-based episodic term in their exploration bonus. As a result, despite their success in relatively simple, noise-free settings, these methods fall short in more realistic scenarios where the state space is vast and prone to noise. To address this limitation, we introduce Exploration via Elliptical Episodic Bonuses (E3B), a new method which extends count-based episodic bonuses to continuous state spaces and encourages an agent to explore states that are diverse under a learned embedding within each episode. The embedding is learned using an inverse dynamics model in order to capture controllable aspects of the environment. Our method sets a new state-of-the-art across 16 challenging tasks from the MiniHack suite, without requiring task-specific inductive b
Exploration via Elliptical Episodic Bonuses Mikael Henaff Meta AI Research mikaelhenaff@meta.com &Roberta Raileanu Meta AI Research raileanu@meta.com aaaaa Minqi Jiang aaaaa University College London aaaaa Meta AI Research aaaaa meta@fb.com &Tim Rocktäschel University College London t.rocktaschel@cs.ucl.ac.uk Abstract In recent years, a number of reinforcement learning (RL) methods have been proposed to explore complex environments which differ across episodes. In this work, we show that the effectiveness of these methods critically relies on a count-based episodic term in their exploration bo
Explore this link on the map →related reading
- Exploration Strategies in Deep Reinforcement Learning | Lil'Loglilianweng.github.io
- Reward Hacking in Reinforcement Learning | Lil'Loglilianweng.github.io
- Exploration for the Efficient Deployment of Reinforcement Learning Agentsopenreview.net
- [1606.01868] Unifying Count-Based Exploration and Intrinsic Motivationar5iv.labs.arxiv.org
- A (Long) Peek into Reinforcement Learning | Lil'Loglilianweng.github.io
- State of RL for reasoning LLMs | A. Weersaweers.de
- Key Papers in Deep RL - Spinning Up documentationspinningup.openai.com
- Reinforcement Learning via Implicit Imitation Guidancearxiv.org
- [1705.05363] Curiosity-driven Exploration by Self-supervised Predictionar5iv.labs.arxiv.org
- The Multi-Armed Bandit Problem and Its Solutions | Lil'Loglilianweng.github.io
- [1901.10995] Go-Explore: a New Approach for Hard-Exploration Problemsar5iv.labs.arxiv.org
- [2201.11861] The Challenges of Exploration for Offline Reinforcement Learningar5iv.labs.arxiv.org