flâneur

Understanding Policy Gradients | John Lambert

johnwlambert.github.io · 4,040 words · saved by 1 readers

A simple, whitespace theme for academics. Based on [*folio](https://github.com/bogoli/-folio) design.

Table of Contents: Policy Gradients Geometric Intuition Math Background Policy Gradient Theorem The Probability of a Trajectory, Given a Policy The REINFORCE Algorithm Baseline Subtraction Vanilla PG: Python Implementation TD Error as Advantage Function Trust Region Policy Optimization (TRPO) Truncated Natural Gradient Policy Algorithm PPO Policy gradients is a reinforcement learning method, where an agent interacts with the world, taking decisions. After taking an action, the world (i.e. environment) changes, and this process repeats over and over and over. We’ll examine the…

saved by

related reading