flâneur — a map of the web's best reading

Is Value Learning Really the Main Bottleneck in Offline RL? | HTML5

ar5iv.labs.arxiv.org · 14,184 words · saved by 1 readers

While imitation learning requires access to high-quality data, offline reinforcement learning (RL) should, in principle, perform similarly or better with substantially lower data quality by using a value function. However, current results indicate that offline RL often performs worse than imitation learning, and it is often unclear what holds back the performance of offline RL. Motivated by this observation, we aim to understand the bottlenecks in current offline RL algorithms. While poor performance of offline RL is typically attributed to an imperfect value function, we ask: is the main bottleneck of offline RL indeed in learning the value function, or something else? To answer this question, we perform a systematic empirical study of (1) value learning, (2) policy extraction, and (3) policy generalization in offline RL problems, analyzing how these components affect performance. We make two surprising observations. First, we find that the choice of a policy extraction algorithm sign

Is Value Learning Really the Main Bottleneck in Offline RL? Seohong Park 1 Kevin Frans 1 Sergey Levine 1 Aviral Kumar 2 1 University of California, Berkeley 2 Google DeepMind seohong@berkeley.edu Abstract While imitation learning requires access to high-quality data, offline reinforcement learning (RL) should, in principle, perform similarly or better with substantially lower data quality by using a value function. However, current results indicate that offline RL often performs worse than imitation learning, and it is often unclear what holds back the performance of offline RL. Motivated by t

Explore this link on the map →

related reading