✳flâneur — a map of the web's best reading
A Graphic Guide to Implementing PPO for Atari Games | by DarylRodrigo | Towards Data Science
towardsdatascience.com · 7,282 words · saved by 1 readers
Written by Daryl and Daniel
A Graphic Guide to Implementing PPO for Atari Games | Towards Data Science A Graphic Guide to Implementing PPO for Atari Games Written by Daryl and Daniel DarylRodrigo Feb 7, 2021 35 min read Share Learnings from a journey to code Proximal Policy Optimisation Written by Daryl and Daniel Image by sergeitokmakov on Pixabay Learning how Proximal Policy Optimisation (PPO) works and writing a functioning version is hard. There are many places where this can go wrong – from misunderstanding the maths and mismatching tensors to having a logical error in the implementation. It took a friend and
Explore this link on the map →saved by
related reading
- The 37 Implementation Details of Proximal Policy Optimization · The ICLR Blog Trackiclr-blog-track.github.io
- A vision researcher’s guide to some RL stuff: PPO & GRPO - Yuge (Jimmy) Shiyugeten.github.io
- Part 1: Key Concepts in RL - Spinning Up documentationspinningup.openai.com
- [1707.06347] Proximal Policy Optimization Algorithmsarxiv.org
- PPO for LLMs: A Guide for Normal Peoplecameronrwolfe.substack.com
- Policy Gradient Algorithms | Lil'Loglilianweng.github.io
- State of RL for reasoning LLMs | A. Weersaweers.de
- RLHF Bookrlhfbook.com
- A (Long) Peek into Reinforcement Learning | Lil'Loglilianweng.github.io
- Part 3: Intro to Policy Optimization - Spinning Up documentationspinningup.openai.com
- Deep Reinforcement Learning: Pong from Pixelskarpathy.github.io
- Proximal Policy Optimization - Spinning Up documentationspinningup.openai.com