A (Long) Peek into Reinforcement Learning | Lil'Log
[Updated on 2020-09-03: Updated the algorithm of SARSA and Q-learning so that the difference is more pronounced. [Updated on 2021-09-19: Thanks to 爱吃猫的鱼, we have this post in Chinese].
Table of Contents What is Reinforcement Learning? Key Concepts Model: Transition and Reward Policy Value Function Optimal Value and Policy Markov Decision Processes Bellman Equations Bellman Expectation Equations Bellman Optimality Equations Common Approaches Dynamic Programming Policy Evaluation Policy Improvement Policy Iteration Monte-Carlo Methods Temporal-Difference Learning Bootstrapping Value Estimation SARSA: On-Policy TD control Q-Learning: Off-policy TD control Deep Q-Network Combining TD and MC Learning Policy Gradient Policy Gradient Theorem REINFORCE Actor-Critic A3C Evolution Str
Explore this link on the map →saved by
- Yufei Xiao
- Chloe Chia
- James Rogers
- Fernando Silva
- Asma Lamgh
- Linda
- Jonah Dykhuizen
- Aaron Pham
- Tony Leapo
- Nathan Chen
related reading
- Policy Gradient Algorithms | Lil'Loglilianweng.github.io
- An Updated Introduction to Reinforcement Learning | Sri's Blogsrianumakonda.com
- Part 1: Key Concepts in RL - Spinning Up documentationspinningup.openai.com
- Deep Reinforcement Learning: Pong from Pixelskarpathy.github.io
- Part 2: Kinds of RL Algorithms - Spinning Up documentationspinningup.openai.com
- Q-learning - Wikipediaen.wikipedia.org
- Reinforcement learning - Wikipediaen.wikipedia.org
- Q-learning is not yet scalableseohong.me
- Deep Q-Networks Explained — LessWronglesswrong.com
- [2005.01643] Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problemsar5iv.labs.arxiv.org
- Key Papers in Deep RL - Spinning Up documentationspinningup.openai.com
- Reinforcement Learning in Newcomblike Problemsproceedings.neurips.cc