✳flâneur — a map of the web's best reading

🗂️ A Taxonomy of Reinforcement Learning Algorithms | Arushi Somani

amks.me · 1,215 words · saved by 1 readers

There is a wide range of different patterns of reinforcement learning concepts, each slightly different from the previous — value iteration, policy iteration, PPO, TRPO, REINFORCE, A2C, A3C… If these are also confusing to you and muddle together into one blob of ideas, this is completely expected! Four decades of the proud tradition of categorization and naming have overloaded terms and letters, and made it impossible for new-comers to understand what’s going on. The goal of this article is not to explain each of these terms— we would be stuck here forever! — but to provide to you a model for the constraints that create the categories that create a bundle of all of these policies. My goal is that, by the end of this, you could look at an algorithm, and be able to place it in a set of categories, understanding trade-offs and what works better for what domains. Below is a flowchart of how to think of RL algorithms. I’ll warn you because you peruse it — it is a simplification. In truth, c

Table of Contents Introduction Categorizations between RL Algorithms Model-Based vs Model-Free Learning Values vs Policies On-Policy vs Off-Policy Pop Quiz: What is AlphaGo? Conclusion Introduction There is a wide range of different patterns of reinforcement learning concepts, each slightly different from the previous — value iteration, policy iteration, PPO, TRPO, REINFORCE, A2C, A3C… If these are also confusing to you and muddle together into one blob of ideas, this is completely expected! Four decades of the proud tradition of categorization and naming have overloaded terms and letters, and

Explore this link on the map →

related reading