Quasimetric RL (QRL)
In goal-reaching reinforcement learning (RL), the optimal value function has a particular geometry, called quasimetric structure (see also these works). This paper introduces Quasimetric Reinforcement Learning (QRL), a new RL method that utilizes quasimetric models to learn optimal value functions. Distinct from prior approaches, the QRL objective is specifically designed for quasimetrics, and provides strong theoretical recovery guarantees. Empirically, we conduct thorough analyses on a discretized MountainCar environment, identifying properties of QRL and its advantages over alternatives. On offline and online goal-reaching benchmarks, QRL also demonstrates improved sample efficiency and performance, across both state-based and image-based observations. ICML 2023. arXiv 2304.01203. Tongzhou Wang, Antonio Torralba, Phillip Isola, Amy Zhang. "Optimal Goal-Reaching Reinforcement Learning via Quasimetric Learning" International Conference on Machine Learning (ICML). 2023.
Quasimetric RL (QRL) Optimal Goal-Reaching Reinforcement Learning via Quasimetric Learning Tongzhou Wang MIT CSAIL Antonio Torralba MIT CSAIL Phillip Isola MIT CSAIL Amy Zhang UT Austin, Meta AI ICML 2023 PMLR Proceedings --> [arXiv] [Code] [Thread🧵] --> Overview of Quasimetric RL (QRL) Quasimetric Geometry + A Novel Objective (Push apart start state and goal while maintaining local distances) = Optimal Value $V^*$ AND High-Performing Goal-Reaching Agents Abstract In goal-reaching reinforcement learning (RL), the optimal value function has a particular geometry, called quasimetric stru
Explore this link on the map →related reading
- A (Long) Peek into Reinforcement Learning | Lil'Loglilianweng.github.io
- Part 1: Key Concepts in RL - Spinning Up documentationspinningup.openai.com
- Q-learning is not yet scalableseohong.me
- State of RL for reasoning LLMs | A. Weersaweers.de
- Part 2: Kinds of RL Algorithms - Spinning Up documentationspinningup.openai.com
- An Updated Introduction to Reinforcement Learning | Sri's Blogsrianumakonda.com
- Accelerating Goal-Conditioned Reinforcement Learning Algorithms and Researcharxiv.org
- Key Papers in Deep RL - Spinning Up documentationspinningup.openai.com
- pistar06.pdfpi.website
- The Promise of Hierarchical Reinforcement Learningthegradient.pub
- [2406.09329] Is Value Learning Really the Main Bottleneck in Offline RL?ar5iv.labs.arxiv.org
- [1905.13341] On Value Functions and the Agent-Environment Boundaryar5iv.labs.arxiv.org