When is Realizability Sufficient for Off-Policy Reinforcement Learning? | HTML5
Understanding when reinforcement learning algorithms can make successful off-policy predictions—and when the may fail to do so–remains an open problem. Typically, model-free algorithms for reinforcement learning are an…
[2211.05311] When is Realizability Sufficient for Off-Policy Reinforcement Learning? When is Realizability Sufficient for Off-Policy Reinforcement Learning? Andrea Zanette Abstract Understanding when reinforcement learning algorithms can make successful off-policy predictions—and when the may fail to do so–remains an open problem. Typically, model-free algorithms for reinforcement learning are analyzed under a condition called Bellman completeness when they operate off-policy with function approximation, unless additional conditions are met. However, Bellman completeness is a requirement that
related reading
- A (Long) Peek into Reinforcement Learning | Lil'Loglilianweng.github.io
- rltheorybook_ABJKS.pdfrltheorybook.github.io
- [2007.08202] Provably Good Batch Reinforcement Learning Without Great Explorationar5iv.labs.arxiv.org
- c21f4ce780c5c9d774f79841b81fdc6d-Paper.pdfproceedings.neurips.cc
- Q-learning is not yet scalableseohong.me
- [2005.01643] Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problemsar5iv.labs.arxiv.org
- ML Mentorship: Some Q/A about RLevjang.com
- [2307.04354] Policy Finetuning in Reinforcement Learning via Design of Experiments using Offline Dataar5iv.labs.arxiv.org
- Value estimation with finite datamcgill.scholaris.ca
- 2208.05129arxiv.org
- Reinforcement Learning in Newcomblike Problemsproceedings.neurips.cc
- [2602.19362] LLMs Can Learn to Reason Via Off-Policy RLarxiv.org