Proximal Policy Optimization (PPO)
huggingface.co · 2,236 words · saved by 1 readers
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
Unit 8, of the Deep Reinforcement Learning Class with Hugging Face 🤗 ⚠️ A new updated version of this article is available here 👉 https://huggingface.co/deep-rl-course/unit1/introduction This article is part of the Deep Reinforcement Learning Class. A free course from beginner to expert. Check the syllabus here. ⚠️ A new updated version of this article is available here 👉 https://huggingface.co/deep-rl-course/unit1/introduction This article is part of the Deep Reinforcement Learning Class. A free course from beginner to expert. Check the syllabus here. In the last Unit, we learned…
related reading
- The 37 Implementation Details of Proximal Policy Optimization · The ICLR Blog Trackiclr-blog-track.github.io
- [1707.06347] Proximal Policy Optimization Algorithmsarxiv.org
- A vision researcher’s guide to some RL stuff: PPO & GRPO - Yuge (Jimmy) Shiyugeten.github.io
- Proximal Policy Optimization Algorithmsalphaxiv.org
- Proximal Policy Optimization - Spinning Up documentationspinningup.openai.com
- A Graphic Guide to Implementing PPO for Atari Games | Towards Data Sciencetowardsdatascience.com
- State of RL for reasoning LLMs | A. Weersaweers.de
- PPO for LLMs: A Guide for Normal Peoplecameronrwolfe.substack.com
- Policy Gradient Algorithms | Lil'Loglilianweng.github.io
- RLHF Bookrlhfbook.com
- A Guide to Reinforcement Learning Post-Training for LLMs: PPO, DPO, GRPO, and Beyondhuggingface.co
- [1707.06347] Proximal Policy Optimization Algorithmsarxiv.org