flâneur — a map of the web's best reading

An Updated Introduction to Reinforcement Learning | Sri's Blog

srianumakonda.com · 9,414 words · saved by 1 readers

A while back I wrote a blog on understanding the fundamentals of RL. I’ve spent the past couple weeks reading through Kevin Murphy’s Reinforcement Learning textbook and Sutton and Barto to review some of my fundamentals. This blog contains some notes to cover topics I haven’t yet talked about in my first attempt at explaining RL! What is Reinforcement Learning? Reinforcement Learning is all about the idea of interacting with your environment to learn good behaviors. Given the full state $s_t$, observation $o_t$, some policy $\pi$, action $a_t = \pi(o_t)$, and reward $r_t$, the goal of an agent is to maximize the sum of its expected rewards:

Table of Contents What is Reinforcement Learning? Markov Decision Processes Bellman Equations Value-based RL Value Iteration Policy Iteration Policy evaluation Policy improvement TD Learning Monte-Carlo Learning TD-$\lambda$ SARSA Q-Learning Double Q-Learning Policy-Based Reinforcement Learning Policy Gradient Theorem REINFORCE Advantage Actor Critic (A2C) Generalize Advantage Estimation (GAE) Deep Deterministic Policy Gradients Twin Delayed DDPG (TD3) Trust Region Policy Optimization (TRPO) Proximal Policy Optimization (PPO) Soft Actor Critic (SAC) Citation References A while back I wrote a b

Explore this link on the map →

saved by

related reading