The 37 Implementation Details of Proximal Policy Optimization · The ICLR Blog Track
Jon is a first-year master’s student who is interested in reinforcement learning (RL). In his eyes, RL seemed fascinating because he could use RL libraries such as Stable-Baselines3 (SB3) to train agents to play all kinds of games. He quickly recognized Proximal Policy Optimization (PPO) as a fast and versatile algorithm and wanted to implement PPO himself as a learning experience. Upon reading the paper, Jon thought to himself, “huh, this is pretty straightforward.” He then opened a code editor and started writing PPO. CartPole-v1 from Gym was his chosen simulation environment, and before long, Jon made PPO work with CartPole-v1. He had a great time and felt motivated to make his PPO work with more interesting environments, such as the Atari games and MuJoCo robotics tasks. “How cool would that be?” he thought. However, he soon struggled. Making PPO work with Atari and MuJoCo seemed more challenging than anticipated. Jon then looked for reference implementations online but was shortly
The 37 Implementation Details of Proximal Policy Optimization · The ICLR Blog Track --> For short-term, peer-sourced tests of time, generalizations, specializations, reproductions, etc.! © 2024. All rights reserved. The ICLR Blog Track | Blog Posts The 37 Implementation Details of Proximal Policy Optimization 25 Mar 2022 | proximal-policy-optimization reproducibility reinforcement-learning implementation-details tutorial Huang, Shengyi; Dossa, Rousslan Fernand Julien; Raffin, Antonin; Kanervisto, Anssi; Wang, Weixun Jon is a first-year master’s student who is interested in reinforc
Explore this link on the map →saved by
- Claire Wang
- Saeejith Nair
- Yixiong Hao
- Lydia Nottingham
- Vincent Cheng
- Yudhister Joel Kumar
- Akira Yoshiyama
- A P
- Evie Hu
related reading
- [1707.06347] Proximal Policy Optimization Algorithmsarxiv.org
- Debugging Reinforcement Learning Systemsandyljones.com
- A vision researcher’s guide to some RL stuff: PPO & GRPO - Yuge (Jimmy) Shiyugeten.github.io
- PPO for LLMs: A Guide for Normal Peoplecameronrwolfe.substack.com
- State of RL for reasoning LLMs | A. Weersaweers.de
- A Graphic Guide to Implementing PPO for Atari Games | Towards Data Sciencetowardsdatascience.com
- [1709.06560] Deep Reinforcement Learning that Mattersarxiv.org
- Policy Gradient Algorithms | Lil'Loglilianweng.github.io
- Proximal Policy Optimization - Spinning Up documentationspinningup.openai.com
- Part 3: Intro to Policy Optimization - Spinning Up documentationspinningup.openai.com
- An Updated Introduction to Reinforcement Learning | Sri's Blogsrianumakonda.com
- Part 1: Key Concepts in RL - Spinning Up documentationspinningup.openai.com