flâneur — a map of the web's best reading

[2006.11266] An operator view of policy gradient methods

ar5iv.labs.arxiv.org · 14,201 words · saved by 1 readers

We cast policy gradient methods as the repeated application of two operators: a policy improvement operator , which maps any policy to a better one , and a projection operator , which finds the best approximation of …

An operator view of policy gradient methods Dibya Ghosh Google Brain &Marlos C. Machado Google Brain &Nicolas Le Roux Google Brain Abstract We cast policy gradient methods as the repeated application of two operators: a policy improvement operator ℐ ℐ {\mathcal{I}} , which maps any policy π 𝜋 \pi to a better one ℐ ​ π ℐ 𝜋 {\mathcal{I}}\pi , and a projection operator 𝒫 𝒫 {\mathcal{P}} , which finds the best approximation of ℐ ​ π ℐ 𝜋 {\mathcal{I}}\pi in the set of realizable policies. We use this framework to introduce operator-based versions of well-known policy gradient methods such as R

Explore this link on the map →

related reading