Is DPO Superior to PPO for LLM Alignment? A Comprehensive Study
This is experimental HTML to improve accessibility. We invite you to report rendering errors. Use Alt+Y to toggle on accessible reporting links and Alt+Shift+Y to toggle off. Learn more about this project and help improve conversions. Reinforcement Learning from Human Feedback (RLHF) is currently the most widely used method to align large language models (LLMs) with human preferences. Existing RLHF methods can be roughly categorized as either reward-based or reward-free. Novel applications such as ChatGPT and Claude leverage reward-based methods that first learn a reward model and apply actor-critic algorithms, such as Proximal Policy Optimization (PPO). However, in academic benchmarks, the state-of-the-art results are often achieved via reward-free methods, such as Direct Preference Optimization (DPO). Is DPO truly superior to PPO? Why does PPO perform poorly on these benchmarks? In this paper, we first conduct both theoretical and empirical studies on the algorithmic properties of D
Is DPO Superior to PPO for LLM Alignment? A Comprehensive Study Shusheng Xu Wei Fu Jiaxuan Gao Wenjie Ye Weilin Liu Zhiyu Mei Guangju Wang Chao Yu Yi Wu Abstract Reinforcement Learning from Human Feedback (RLHF) is currently the most widely used method to align large language models (LLMs) with human preferences. Existing RLHF methods can be roughly categorized as either reward-based or reward-free . Novel applications such as ChatGPT and Claude leverage reward-based methods that first learn a reward model and apply actor-critic algorithms, such as Proximal Policy Optimization (PPO). However,
Explore this link on the map →related reading
- [2305.18290] Direct Preference Optimization: Your Language Model is Secretly a Reward Modelarxiv.org
- State of RL for reasoning LLMs | A. Weersaweers.de
- A vision researcher’s guide to some RL stuff: PPO & GRPO - Yuge (Jimmy) Shiyugeten.github.io
- A Guide to Reinforcement Learning Post-Training for LLMs: PPO, DPO, GRPO, and Beyondhuggingface.co
- Rethinking the Role of PPO in RLHF – The Berkeley Artificial Intelligence Research Blogbair.berkeley.edu
- Understanding the Performance Gap in Preference Learning: A Dichotomy of RLHF and DPOarxiv.org
- RLHF & Post-Training Course by Nathan Lambertrlhfbook.com
- PPO for LLMs: A Guide for Normal Peoplecameronrwolfe.substack.com
- [2402.14740] Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMsarxiv.org
- RLHF: Reinforcement Learning from Human Feedbackhuyenchip.com
- The State of Reinforcement Learning for LLM Reasoningmagazine.sebastianraschka.com
- [2605.03327] DGPO: Distribution Guided Policy Optimization for Fine Grained Credit Assignmentarxiv.org