🗂️ A Taxonomy of Reinforcement Learning Algorithms | Arushi Somani
There is a wide range of different patterns of reinforcement learning concepts, each slightly different from the previous — value iteration, policy iteration, PPO, TRPO, REINFORCE, A2C, A3C… If these are also confusing to you and muddle together into one blob of ideas, this is completely expected! Four decades of the proud tradition of categorization and naming have overloaded terms and letters, and made it impossible for new-comers to understand what’s going on. The goal of this article is not to explain each of these terms— we would be stuck here forever! — but to provide to you a model for the constraints that create the categories that create a bundle of all of these policies. My goal is that, by the end of this, you could look at an algorithm, and be able to place it in a set of categories, understanding trade-offs and what works better for what domains. Below is a flowchart of how to think of RL algorithms. I’ll warn you because you peruse it — it is a simplification. In truth, c
Table of Contents Introduction Categorizations between RL Algorithms Model-Based vs Model-Free Learning Values vs Policies On-Policy vs Off-Policy Pop Quiz: What is AlphaGo? Conclusion Introduction There is a wide range of different patterns of reinforcement learning concepts, each slightly different from the previous — value iteration, policy iteration, PPO, TRPO, REINFORCE, A2C, A3C… If these are also confusing to you and muddle together into one blob of ideas, this is completely expected! Four decades of the proud tradition of categorization and naming have overloaded terms and letters, and
Explore this link on the map →related reading
- Part 2: Kinds of RL Algorithms - Spinning Up documentationspinningup.openai.com
- A (Long) Peek into Reinforcement Learning | Lil'Loglilianweng.github.io
- Part 1: Key Concepts in RL - Spinning Up documentationspinningup.openai.com
- Model-free (reinforcement learning) - Wikipediaen.wikipedia.org
- Key Papers in Deep RL - Spinning Up documentationspinningup.openai.com
- An Updated Introduction to Reinforcement Learning | Sri's Blogsrianumakonda.com
- A vision researcher’s guide to some RL stuff: PPO & GRPO - Yuge (Jimmy) Shiyugeten.github.io
- Reinforcement learning - Wikipediaen.wikipedia.org
- RLHF Bookrlhfbook.com
- Policy Gradient Algorithms | Lil'Loglilianweng.github.io
- [2005.01643] Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problemsar5iv.labs.arxiv.org
- Q-learning - Wikipediaen.wikipedia.org