When is Realizability Sufficient for Off-Policy Reinforcement Learning? | HTML5
Understanding when reinforcement learning algorithms can make successful off-policy predictions—and when the may fail to do so–remains an open problem. Typically, model-free algorithms for reinforcement learning are an…
[2211.05311] When is Realizability Sufficient for Off-Policy Reinforcement Learning? When is Realizability Sufficient for Off-Policy Reinforcement Learning? Andrea Zanette Abstract Understanding when reinforcement learning algorithms can make successful off-policy predictions—and when the may fail to do so–remains an open problem. Typically, model-free algorithms for reinforcement learning are analyzed under a condition called Bellman completeness when they operate off-policy with function approximation, unless additional conditions are met. However, Bellman completeness is a requirement that
Explore this link on the map →related reading
- A (Long) Peek into Reinforcement Learning | Lil'Loglilianweng.github.io
- Q-learning is not yet scalableseohong.me
- [2007.08202] Provably Good Batch Reinforcement Learning Without Great Explorationar5iv.labs.arxiv.org
- [2005.01643] Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problemsar5iv.labs.arxiv.org
- [2307.04354] Policy Finetuning in Reinforcement Learning via Design of Experiments using Offline Dataar5iv.labs.arxiv.org
- pdfopenreview.net
- [1905.13341] On Value Functions and the Agent-Environment Boundaryar5iv.labs.arxiv.org
- State of RL for reasoning LLMs | A. Weersaweers.de
- Exploration for the Efficient Deployment of Reinforcement Learning Agentsopenreview.net
- Reinforcement Learning in Newcomblike Problemsproceedings.neurips.cc
- An Updated Introduction to Reinforcement Learning | Sri's Blogsrianumakonda.com
- Overcoming the Sim-to-Real Gap: Leveraging Simulation to Learn to Explore for Real-World RLarxiv.org