✳flâneur — a map of the web's best reading
Rethinking the Role of PPO in RLHF – The Berkeley Artificial Intelligence Research Blog
bair.berkeley.edu · 1,695 words · saved by 1 readers
The BAIR Blog
Rethinking the Role of PPO in RLHF TL;DR : In RLHF, there’s tension between the reward learning phase, which uses human preference in the form of comparisons, and the RL fine-tuning phase, which optimizes a single, non-comparative reward. What if we performed RL in a comparative way? Figure 1: This diagram illustrates the difference between reinforcement learning from absolute feedback and relative feedback. By incorporating a new component - pairwise policy gradient, we can unify the reward modeling stage and RL stage, enabling direct updates based on pairwise responses. Large Language Models
Explore this link on the map →related reading
- A vision researcher’s guide to some RL stuff: PPO & GRPO - Yuge (Jimmy) Shiyugeten.github.io
- State of RL for reasoning LLMs | A. Weersaweers.de
- PPO for LLMs: A Guide for Normal Peoplecameronrwolfe.substack.com
- A Guide to Reinforcement Learning Post-Training for LLMs: PPO, DPO, GRPO, and Beyondhuggingface.co
- [2305.18290] Direct Preference Optimization: Your Language Model is Secretly a Reward Modelarxiv.org
- RLHF: Reinforcement Learning from Human Feedbackhuyenchip.com
- [2402.14740] Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMsarxiv.org
- RLHF & Post-Training Course by Nathan Lambertrlhfbook.com
- The N Implementation Details of RLHF with PPO | ICLR Blogposts 2024iclr-blogposts.github.io
- RLHF Bookrlhfbook.com
- Is DPO Superior to PPO for LLM Alignment? A Comprehensive Studyarxiv.org
- rlhfbook.com/book.pdfrlhfbook.com