flâneur — a map of the web's best reading

A (Long) Peek into Reinforcement Learning | Lil'Log

lilianweng.github.io · 6,575 words · saved by 10 readers

[Updated on 2020-09-03: Updated the algorithm of SARSA and Q-learning so that the difference is more pronounced. [Updated on 2021-09-19: Thanks to 爱吃猫的鱼, we have this post in Chinese].

Table of Contents What is Reinforcement Learning? Key Concepts Model: Transition and Reward Policy Value Function Optimal Value and Policy Markov Decision Processes Bellman Equations Bellman Expectation Equations Bellman Optimality Equations Common Approaches Dynamic Programming Policy Evaluation Policy Improvement Policy Iteration Monte-Carlo Methods Temporal-Difference Learning Bootstrapping Value Estimation SARSA: On-Policy TD control Q-Learning: Off-policy TD control Deep Q-Network Combining TD and MC Learning Policy Gradient Policy Gradient Theorem REINFORCE Actor-Critic A3C Evolution Str

Explore this link on the map →

saved by

related reading