[2109.08603] Is Curiosity All You Need? On the Utility of Emergent Behaviours from Curious Exploration
Curiosity-based reward schemes can present powerful exploration mechanisms which facilitate the discovery of solutions for complex, sparse or long-horizon tasks. However, as the agent learns to reach previously unexplored spaces and the objective adapts to reward new areas, many behaviours emerge only to disappear due to being overwritten by the constantly shifting objective. We argue that merely using curiosity for fast environment exploration or as a bonus reward for a specific task does not harness the full potential of this technique and misses useful skills. Instead, we propose to shift the focus towards retaining the behaviours which emerge during curiosity-based learning. We posit that these self-discovered behaviours serve as valuable skills in an agent's repertoire to solve related tasks. Our experiments demonstrate the continuous shift in behaviour throughout training and the benefits of a simple policy snapshot method to reuse discovered behaviour for transfer tasks.
View PDF HTML (experimental) Abstract:Curiosity-based reward schemes can present powerful exploration mechanisms which facilitate the discovery of solutions for complex, sparse or long-horizon tasks. However, as the agent learns to reach previously unexplored spaces and the objective adapts to reward new areas, many behaviours emerge only to disappear due to being overwritten by the constantly shifting objective. We argue that merely using curiosity for fast environment exploration or as a bonus reward for a specific task does not harness the full potential of this technique and misses…
saved by
related reading
- Multi-task curriculum learning in a complex, visual,hard-exploration domain: Minecraftarxiv.org
- [1705.05363] Curiosity-driven Exploration by Self-supervised Predictionar5iv.labs.arxiv.org
- Complex behavior from intrinsic motivation to occupy future action-state path spacenature.com
- Reward is not the optimization target — LessWronglesswrong.com
- [2507.09041] Behavioral Exploration: Learning to Explore via In-Context Adaptationar5iv.labs.arxiv.org
- go-explore-nature.pdfadrien.ecoffet.com
- Voyager | An Open-Ended Embodied Agent with Large Language Modelsvoyager.minedojo.org
- 0812.4360arxiv.org
- Exploration Strategies in Deep Reinforcement Learning | Lil'Loglilianweng.github.io
- Learning Beyond Gradientstrinkle23897.github.io
- Reward Is Not the Optimization Targetturntrout.com
- [2210.05805] Exploration via Elliptical Episodic Bonusesar5iv.labs.arxiv.org