A vision researcher’s guide to some RL stuff: PPO & GRPO - Yuge (Jimmy) Shi
It has been a while since I last wrote a blog post. Life has been hectic since I started work, and the machine learning world is also not what it was since I graduated in early 2023. Your average parents having LLM apps installed on their phones is already yesterday’s news – I took two weeks off work to spend Lunar New Year in China, which only serves to give me plenty of time to scroll on twitter and witness DeepSeek’s (quite well-deserved) hype peak on Lunar New Year’s eve while getting completely overwhelmed. So this feels like a good time to read, learn, do some basic maths, and write some stuff down again. This is a deep dive into Proximal Policy Optimization (PPO), which is one of the most popular algorithm used in RLHF for LLMs, as well as Group Relative Policy Optimization (GRPO) proposed by the DeepSeek folks, and there’s also a quick summary of the tricks I find impressive in the DeepSeek R1 tech report in the end. This is all done by someone who’s mostly worked on vision and
First up, some rambles as usual. It has been a while since I last wrote a blog post. Life has been hectic since I started work, and the machine learning world is also not what it was since I graduated in early 2023. Your average parents having LLM apps installed on their phones is already yesterday’s news – I took two weeks off work to spend Lunar New Year in China, which only serves to give me plenty of time to scroll on twitter and witness DeepSeek’s (quite well-deserved) hype peak on Lunar New Year’s eve while getting completely overwhelmed. So this feels like a good time to read, learn, do
Explore this link on the map →saved by
- Ratan Kaliani
- Lydia Nottingham
- Julian H
- Jackson Mowatt Gok
- Ishan Mukherjee
- Ishaan Panigrahi
- Jeff Brown
- Kevin Zhang
related reading
- State of RL for reasoning LLMs | A. Weersaweers.de
- A Guide to Reinforcement Learning Post-Training for LLMs: PPO, DPO, GRPO, and Beyondhuggingface.co
- From REINFORCE to Dr. GRPOlancelqf.github.io
- Why GRPO is Important and How it Worksghost.oxen.ai
- PPO for LLMs: A Guide for Normal Peoplecameronrwolfe.substack.com
- The State of Reinforcement Learning for LLM Reasoningmagazine.sebastianraschka.com
- RL Scaling Laws for LLMs - by Cameron R. Wolfe, Ph.D.cameronrwolfe.substack.com
- RLHF Bookrlhfbook.com
- The 37 Implementation Details of Proximal Policy Optimization · The ICLR Blog Trackiclr-blog-track.github.io
- Xiuyu Li on X: "RL Interview Questions 2026" / Xx.com
- DeepSeek-R1arxiv.org
- rlhfbook.com/book.pdfrlhfbook.com