Actor-Critic Algorithms
We propose and analyze a class of actor-critic algorithms for simulation-based optimization of a Markov decision process over a parameterized family of randomized stationary policies. These are two-time-scale algorithms in which the critic uses TD learning with a linear approximation architecture and the actor is updated in an approximate gradient direction based on information pro(cid:173) vided by the critic. We show that the features for the critic should span a subspace prescribed by the choice of parameterization of the actor. We conclude by discussing convergence properties and some open problems.
Vijay R. Konda, John N. Tsitsiklis Advances in Neural Information Processing Systems 12 (NIPS 1999) Abstract We propose and analyze a class of actor-critic algorithms for simulation-based optimization of a Markov decision process over a parameterized family of randomized stationary policies. These are two-time-scale algorithms in which the critic uses TD learning with a linear approximation architecture and the actor is updated in an approximate gradient direction based on information pro(cid:173) vided by the critic. We show that the features for the critic should span a subspace…
saved by
related reading
- A (Long) Peek into Reinforcement Learning | Lil'Loglilianweng.github.io
- Policy Gradient Algorithms | Lil'Loglilianweng.github.io
- Soft Actor-Critic - Spinning Up documentationspinningup.openai.com
- Extended Data Fig. 5: Training progress. | Naturenature.com
- RL ALGOk-a.in
- Advantage Actor Critic (A2C)huggingface.co
- [1707.06347] Proximal Policy Optimization Algorithmsarxiv.org
- RLAlgsInMDPs.pdfsites.ualberta.ca
- NeurIPS-2021-understanding-end-to-end-model-based-reinforcement-learning-methods-as-implicit-parameterization-Supplemental.pdflis.csail.mit.edu
- The idea behind Actor-Critics and how A2C and A3C improve them | AI Summertheaisummer.com
- [2506.22401] Exploration from a Primal-Dual Lens: Value-Incentivized Actor-Critic Methods for Sample-Efficient Online RLarxiv.org
- RL_Notes__final_.pdfjubayer-ibn-hamid.github.io