Exploration via Elliptical Episodic Bonuses | HTML5
In recent years, a number of reinforcement learning (RL) methods have been proposed to explore complex environments which differ across episodes. In this work, we show that the effectiveness of these methods critically relies on a count-based episodic term in their exploration bonus. As a result, despite their success in relatively simple, noise-free settings, these methods fall short in more realistic scenarios where the state space is vast and prone to noise. To address this limitation, we introduce Exploration via Elliptical Episodic Bonuses (E3B), a new method which extends count-based episodic bonuses to continuous state spaces and encourages an agent to explore states that are diverse under a learned embedding within each episode. The embedding is learned using an inverse dynamics model in order to capture controllable aspects of the environment. Our method sets a new state-of-the-art across 16 challenging tasks from the MiniHack suite, without requiring task-specific inductive b
Exploration via Elliptical Episodic Bonuses Mikael Henaff Meta AI Research mikaelhenaff@meta.com &Roberta Raileanu Meta AI Research raileanu@meta.com aaaaa Minqi Jiang aaaaa University College London aaaaa Meta AI Research aaaaa meta@fb.com &Tim Rocktäschel University College London t.rocktaschel@cs.ucl.ac.uk Abstract In recent years, a number of reinforcement learning (RL) methods have been proposed to explore complex environments which differ across episodes. In this work, we show that the effectiveness of these methods critically relies on a count-based episodic term in their exploration bo
related reading
- Exploration Strategies in Deep Reinforcement Learning | Lil'Loglilianweng.github.io
- go-explore-nature.pdfadrien.ecoffet.com
- Reward Hacking in Reinforcement Learning | Lil'Loglilianweng.github.io
- Exploration for the Efficient Deployment of Reinforcement Learning Agentsopenreview.net
- [1606.01868] Unifying Count-Based Exploration and Intrinsic Motivationar5iv.labs.arxiv.org
- Multi-task curriculum learning in a complex, visual,hard-exploration domain: Minecraftarxiv.org
- State of RL for reasoning LLMs | A. Weersaweers.de
- Key Papers in Deep RL - Spinning Up documentationspinningup.openai.com
- Is Curiosity All You Need? On the Utility of Emergent Behaviours from Curious Explorationarxiv.org
- Reinforcement Learning via Implicit Imitation Guidancearxiv.org
- A Free Lunch from the Noise:Provable and Practical Exploration for Representation Learningarxiv.org
- [2507.13181] Spectral Bellman Method: Unifying Representation and Exploration in RLarxiv.org