The Challenges of Exploration for Offline Reinforcement Learning | HTML5
Offline Reinforcement Learning (ORL) enables us to separately study the two interlinked processes of reinforcement learning: collecting informative experience and inferring optimal behaviour. The second step has been widely studied in the offline setting, but just as critical to data-efficient RL is the collection of informative data. The task-agnostic setting for data collection, where the task is not known a priori, is of particular interest due to the possibility of collecting a single dataset and using it to solve several downstream tasks as they arise. We investigate this setting via curiosity-based intrinsic motivation, a family of exploration methods which encourage the agent to explore those states or transitions it has not yet learned to model. With Explore2Offline, we propose to evaluate the quality of collected data by transferring the collected data and inferring policies with reward relabelling and standard offline RL algorithms. We evaluate a wide variety of data collecti
The Challenges of Exploration for Offline Reinforcement Learning Nathan Lambert Markus Wulfmeier William Whitney Arunkumar Byravan Michael Bloesch Vibhavari Dasagi Tim Hertweck Martin Riedmiller Abstract Offline Reinforcement Learning (ORL) enables us to separately study the two interlinked processes of reinforcement learning: collecting informative experience and inferring optimal behaviour. The second step has been widely studied in the offline setting, but just as critical to data-efficient RL is the collection of informative data. The task-agnostic setting for data collection, where the ta
related reading
- Exploration for the Efficient Deployment of Reinforcement Learning Agentsopenreview.net
- [2103.06326] S4RL: Surprisingly Simple Self-Supervision for Offline Reinforcement Learningarxiv.org
- Should I Use Offline RL or Imitation Learning? – The Berkeley Artificial Intelligence Research Blogbair.berkeley.edu
- 2109.10813.pdfarxiv.org
- Key Papers in Deep RL - Spinning Up documentationspinningup.openai.com
- Reinforcement Learning via Implicit Imitation Guidancearxiv.org
- [2204.05618] When Should We Prefer Offline Reinforcement Learning Over Behavioral Cloning?ar5iv.labs.arxiv.org
- [2307.04354] Policy Finetuning in Reinforcement Learning via Design of Experiments using Offline Dataar5iv.labs.arxiv.org
- [2005.01643] Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problemsar5iv.labs.arxiv.org
- [2406.09329] Is Value Learning Really the Main Bottleneck in Offline RL?ar5iv.labs.arxiv.org
- [2006.03647] Deployment-Efficient Reinforcement Learning via Model-Based Offline Optimizationar5iv.labs.arxiv.org
- Value estimation with finite datamcgill.scholaris.ca