go-explore-nature.pdf
adrien.ecoffet.com · 6,528 words · saved by 1 readers
N/A
First return, then explore Adrien Ecoffet∗1,2 , Joost Huizinga∗1,2 , Joel Lehman1,2 , Kenneth O. Stanley1,2 & Jeff Clune1,2 1 Uber AI Labs, San Francisco, CA, USA 2 OpenAI, San Francisco, CA, USA ∗ These authors contributed equally to this work Correspondence should be addressed to Adrien Ecoffet (email: adrienecoffet@gmail.com), Joost Huizinga (email: joost.hui@gmail.com), and Jeff Clune (email: jclune@gmail.com). Please cite as: Ecoffet, A., Huizinga, J., Lehman, J., Stanley, K.O. and Clune, J. First return, then explore. Nature 590, 580–586 (2021).…
saved by
related reading
- [1901.10995] Go-Explore: a New Approach for Hard-Exploration Problemsar5iv.labs.arxiv.org
- Exploration Strategies in Deep Reinforcement Learning | Lil'Loglilianweng.github.io
- A (Long) Peek into Reinforcement Learning | Lil'Loglilianweng.github.io
- Exploration for the Efficient Deployment of Reinforcement Learning Agentsopenreview.net
- Key Papers in Deep RL - Spinning Up documentationspinningup.openai.com
- Reinforcement Learning via Implicit Imitation Guidancearxiv.org
- Deep Reinforcement Learning Doesn't Work Yetalexirpan.com
- [2507.09041] Behavioral Exploration: Learning to Explore via In-Context Adaptationar5iv.labs.arxiv.org
- Interactive Visualization of RL Algorithms for LLM Trainingzcy233035.github.io
- Learning Beyond Gradientstrinkle23897.github.io
- The Multi-Armed Bandit Problem and Its Solutions | Lil'Loglilianweng.github.io
- [2412.14135] Scaling of Search and Learning: A Roadmap to Reproduce o1 from Reinforcement Learning Perspectivearxiv.org