Towards Deployment-Efficient Reinforcement Learning: Lower Bound and Optimality | HTML5
Deployment efficiency is an important criterion for many real-world applications of reinforcement learning (RL). Despite the community’s increasing interest, there lacks a formal theoretical formulation for the problem. In this paper, we propose such a formulation for deployment-efficient RL (DE-RL) from an “optimization with constraints” perspective: we are interested in exploring an MDP and obtaining a near-optimal policy within minimal deployment complexity, whereas in each deployment the policy can sample a large batch of data. Using finite-horizon linear MDPs as a concrete structural model, we reveal the fundamental limit in achieving deployment efficiency by establishing information-theoretic lower bounds, and provide algorithms that achieve the optimal deployment efficiency. Moreover, our formulation for DE-RL is flexible and can serve as a building block for other practically relevant settings; we give “Safe DE-RL” and “Sample-Efficient DE-RL” as two examples, which may be wort
[2202.06450] Towards Deployment-Efficient Reinforcement Learning: Lower Bound and Optimality Towards Deployment-Efficient Reinforcement Learning: Lower Bound and Optimality Jiawei Huang † † ~{}\dagger , Jinglin Chen † † \dagger , Li Zhao ‡ ‡ \ddagger , Tao Qin ‡ ‡ \ddagger , Nan Jiang † † \dagger , Tie-Yan Liu ‡ ‡ \ddagger † † \dagger Department of Computer Science, University of Illinois at Urbana-Champaign {jiaweih, jinglinc, nanjiang}@illinois.edu ‡ ‡ \ddagger Microsoft Research Asia {lizo, taoqin, tyliu}@microsoft.com Work done during the internship at Microsoft Research Asia. Abstract Dep
Explore this link on the map →related reading
- Exploration for the Efficient Deployment of Reinforcement Learning Agentsopenreview.net
- [2006.03647] Deployment-Efficient Reinforcement Learning via Model-Based Offline Optimizationar5iv.labs.arxiv.org
- Part 1: Key Concepts in RL - Spinning Up documentationspinningup.openai.com
- A (Long) Peek into Reinforcement Learning | Lil'Loglilianweng.github.io
- Key Papers in Deep RL - Spinning Up documentationspinningup.openai.com
- State of RL for reasoning LLMs | A. Weersaweers.de
- IsoCompute Playbook: Optimally Scaling Sampling Compute for RL Training of LLMscompute-optimal-rl-llm-scaling.github.io
- Reinforcement Learning in Newcomblike Problemsproceedings.neurips.cc
- [2307.04354] Policy Finetuning in Reinforcement Learning via Design of Experiments using Offline Dataar5iv.labs.arxiv.org
- [1709.06560] Deep Reinforcement Learning that Mattersarxiv.org
- [2007.08202] Provably Good Batch Reinforcement Learning Without Great Explorationar5iv.labs.arxiv.org
- The Promise of Hierarchical Reinforcement Learningthegradient.pub