On Value Functions and the Agent-Environment Boundary | HTML5
When function approximation is deployed in reinforcement learning (RL), the same problem may be formulated in different ways, often by treating a pre-processing step as a part of the environment or as part of the agent. As a consequence, fundamental concepts in RL, such as (optimal) value functions, are not uniquely defined as they depend on where we draw this agent-environment boundary. This causes further problems in theoretical analyses that provide optimality guarantees, as the same analysis may yield different bounds in equivalent formulations of the same problem. We address this issue via a simple and novel boundary-invariant analysis of Fitted Q-Iteration, a representative RL algorithm, where the assumptions and the guarantees are invariant to the choice of boundary. We also discuss closely related issues on state resetting, deterministic vs stochastic systems, imitation learning, and the verifiability of theoretical assumptions from data. A large part of RL theory—including tha
On Value Functions and the Agent-Environment Boundary Nan Jiang Abstract When function approximation is deployed in reinforcement learning (RL), the same problem may be formulated in different ways, often by treating a pre-processing step as a part of the environment or as part of the agent. As a consequence, fundamental concepts in RL, such as (optimal) value functions, are not uniquely defined as they depend on where we draw this agent-environment boundary . This causes further problems in theoretical analyses that provide optimality guarantees, as the same analysis may yield different bound
Explore this link on the map →related reading
- A (Long) Peek into Reinforcement Learning | Lil'Loglilianweng.github.io
- Part 1: Key Concepts in RL - Spinning Up documentationspinningup.openai.com
- [2007.08202] Provably Good Batch Reinforcement Learning Without Great Explorationar5iv.labs.arxiv.org
- An Updated Introduction to Reinforcement Learning | Sri's Blogsrianumakonda.com
- Reinforcement Learning in Newcomblike Problemsproceedings.neurips.cc
- [2211.05311] When is Realizability Sufficient for Off-Policy Reinforcement Learning?ar5iv.labs.arxiv.org
- Reward is not the optimization target — LessWronglesswrong.com
- Learn Reinforcement Learning (2) - DQN · greentec's bloggreentec.github.io
- [2506.22401] Exploration from a Primal-Dual Lens: Value-Incentivized Actor-Critic Methods for Sample-Efficient Online RLarxiv.org
- Exploration for the Efficient Deployment of Reinforcement Learning Agentsopenreview.net
- Reinforcement learning - Wikipediaen.wikipedia.org
- [1606.05312] Successor Features for Transfer in Reinforcement Learningarxiv.org