Is Value Learning Really the Main Bottleneck in Offline RL? | HTML5
While imitation learning requires access to high-quality data, offline reinforcement learning (RL) should, in principle, perform similarly or better with substantially lower data quality by using a value function. However, current results indicate that offline RL often performs worse than imitation learning, and it is often unclear what holds back the performance of offline RL. Motivated by this observation, we aim to understand the bottlenecks in current offline RL algorithms. While poor performance of offline RL is typically attributed to an imperfect value function, we ask: is the main bottleneck of offline RL indeed in learning the value function, or something else? To answer this question, we perform a systematic empirical study of (1) value learning, (2) policy extraction, and (3) policy generalization in offline RL problems, analyzing how these components affect performance. We make two surprising observations. First, we find that the choice of a policy extraction algorithm sign
Is Value Learning Really the Main Bottleneck in Offline RL? Seohong Park 1 Kevin Frans 1 Sergey Levine 1 Aviral Kumar 2 1 University of California, Berkeley 2 Google DeepMind seohong@berkeley.edu Abstract While imitation learning requires access to high-quality data, offline reinforcement learning (RL) should, in principle, perform similarly or better with substantially lower data quality by using a value function. However, current results indicate that offline RL often performs worse than imitation learning, and it is often unclear what holds back the performance of offline RL. Motivated by t
Explore this link on the map →related reading
- Should I Use Offline RL or Imitation Learning? – The Berkeley Artificial Intelligence Research Blogbair.berkeley.edu
- [2204.05618] When Should We Prefer Offline Reinforcement Learning Over Behavioral Cloning?ar5iv.labs.arxiv.org
- A (Long) Peek into Reinforcement Learning | Lil'Loglilianweng.github.io
- pistar06.pdfpi.website
- [2005.01643] Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problemsar5iv.labs.arxiv.org
- Q-learning is not yet scalableseohong.me
- Exploration for the Efficient Deployment of Reinforcement Learning Agentsopenreview.net
- Part 1: Key Concepts in RL - Spinning Up documentationspinningup.openai.com
- [2201.11861] The Challenges of Exploration for Offline Reinforcement Learningar5iv.labs.arxiv.org
- π*0.6: a VLA That Learns From Experiencephysicalintelligence.company
- Debugging Reinforcement Learning Systemsandyljones.com
- State of Robot Learning, December 2025vedder.io