Reinforcement Learning | RLHF and Post-Training Book by Nathan Lambert
rlhfbook.com · 15,424 words · saved by 1 readers
Policy gradient methods for RLHF and LLM post-training, including PPO, REINFORCE, RLOO, GRPO, and implementation details.
--> RLHF Book --> Reinforcement Learning | RLHF and Post-Training Book by Nathan Lambert Reinforcement Learning from Human Feedback A short introduction to RLHF and post-training focused on language models. Nathan Lambert Lecture 3: Understanding Policy Gradient Algorithms for RL on LLMs Lecture 4: Implementing RL Algorithms for LLMs Reinforcement Learning In the RLHF process, the reinforcement learning algorithm slowly updates the model’s weights with respect to feedback from a reward model. The policy – the model being trained – generates completions to prompts in the training set, then the
saved by
related reading
- Why GRPO is Important and How it Worksghost.oxen.ai
- A vision researcher’s guide to some RL stuff: PPO & GRPO - Yuge (Jimmy) Shiyugeten.github.io
- State of RL for reasoning LLMs | A. Weersaweers.de
- From REINFORCE to Dr. GRPOlancelqf.github.io
- A Guide to Reinforcement Learning Post-Training for LLMs: PPO, DPO, GRPO, and Beyondhuggingface.co
- PPO for LLMs: A Guide for Normal Peoplecameronrwolfe.substack.com
- Harsh Bhatt (@harshbhatt7585) on Xx.com
- GRPO++: Tricks for Making RL Actually Workcameronrwolfe.substack.com
- RLHF | John Lambertjohnwlambert.github.io
- Interactive Visualization of RL Algorithms for LLM Trainingzcy233035.github.io
- Understanding Policy Gradients | John Lambertjohnwlambert.github.io
- RLHF & Post-Training Course by Nathan Lambertrlhfbook.com