[1805.00909] Reinforcement Learning and Control as Probabilistic Inference: Tutorial and Review
arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs. arXiv Operational Status Get status notifications via email or slack
View PDF HTML (experimental) Abstract:The framework of reinforcement learning or optimal control provides a mathematical formalization of intelligent decision making that is powerful and broadly applicable. While the general form of the reinforcement learning problem enables effective reasoning about uncertainty, the connection between reinforcement learning and inference in probabilistic models is not immediately obvious. However, such a connection has considerable value when it comes to algorithm design: formalizing a problem as probabilistic inference in principle allows us to bring to…
saved by
related reading
- What is AIXI?jan.leike.name
- Reinforcement Learning in Newcomblike Problemsproceedings.neurips.cc
- A (Long) Peek into Reinforcement Learning | Lil'Loglilianweng.github.io
- Part 1: Key Concepts in RL - Spinning Up documentationspinningup.openai.com
- [1805.12114] Deep Reinforcement Learning in a Handful of Trials using Probabilistic Dynamics Modelsar5iv.labs.arxiv.org
- Reinforcement learning - Wikipediaen.wikipedia.org
- SuttonBartoIPRLBook2ndEd.pdfweb.stanford.edu
- course_stat_rl.pdfmit.edu
- [2006.10701] Deep Reinforcement Learning amidst Lifelong Non-Stationarityar5iv.labs.arxiv.org
- [2510.13651] What is the objective of reasoning with reinforcement learning?arxiv.org
- Contentsermongroup.github.io
- NeurIPS-2021-understanding-end-to-end-model-based-reinforcement-learning-methods-as-implicit-parameterization-Supplemental.pdflis.csail.mit.edu