flâneur — a map of the web's best reading

The 37 Implementation Details of Proximal Policy Optimization · The ICLR Blog Track

iclr-blog-track.github.io · 11,878 words · saved by 9 readers

Jon is a first-year master’s student who is interested in reinforcement learning (RL). In his eyes, RL seemed fascinating because he could use RL libraries such as Stable-Baselines3 (SB3) to train agents to play all kinds of games. He quickly recognized Proximal Policy Optimization (PPO) as a fast and versatile algorithm and wanted to implement PPO himself as a learning experience. Upon reading the paper, Jon thought to himself, “huh, this is pretty straightforward.” He then opened a code editor and started writing PPO. CartPole-v1 from Gym was his chosen simulation environment, and before long, Jon made PPO work with CartPole-v1. He had a great time and felt motivated to make his PPO work with more interesting environments, such as the Atari games and MuJoCo robotics tasks. “How cool would that be?” he thought. However, he soon struggled. Making PPO work with Atari and MuJoCo seemed more challenging than anticipated. Jon then looked for reference implementations online but was shortly

The 37 Implementation Details of Proximal Policy Optimization · The ICLR Blog Track --> For short-term, peer-sourced tests of time, generalizations, specializations, reproductions, etc.! © 2024. All rights reserved. The ICLR Blog Track | Blog Posts The 37 Implementation Details of Proximal Policy Optimization 25 Mar 2022 | proximal-policy-optimization reproducibility reinforcement-learning implementation-details tutorial Huang, Shengyi; Dossa, Rousslan Fernand Julien; Raffin, Antonin; Kanervisto, Anssi; Wang, Weixun Jon is a first-year master’s student who is interested in reinforc

Explore this link on the map →

saved by

related reading