Understanding the Performance Gap in Preference Learning: A Dichotomy of RLHF and DPO
This is experimental HTML to improve accessibility. We invite you to report rendering errors. Use Alt+Y to toggle on accessible reporting links and Alt+Shift+Y to toggle off. Learn more about this project and help improve conversions.
Understanding the Performance Gap in Preference Learning: A Dichotomy of RLHF and DPO Ruizhe Shi Tsinghua University srz21@mails.tsinghua.edu.cn &Minhak Song ∗ * KAIST minhaksong@kaist.ac.kr Runlong Zhou University of Washington vectorzh@cs.washington.edu &Zihan Zhang University of Washington zihanz46@uw.edu Maryam Fazel University of Washington mfazel@uw.edu & Simon S. Du University of Washington ssdu@cs.washington.edu Equal contribution.Work done while Minhak Song was visiting the University of Washington. Abstract We present a fine-grained theoretical analysis of the performance gap between
Explore this link on the map →related reading
- [2305.18290] Direct Preference Optimization: Your Language Model is Secretly a Reward Modelarxiv.org
- Is DPO Superior to PPO for LLM Alignment? A Comprehensive Studyarxiv.org
- RLHF & Post-Training Course by Nathan Lambertrlhfbook.com
- State of RL for reasoning LLMs | A. Weersaweers.de
- A vision researcher’s guide to some RL stuff: PPO & GRPO - Yuge (Jimmy) Shiyugeten.github.io
- A Guide to Reinforcement Learning Post-Training for LLMs: PPO, DPO, GRPO, and Beyondhuggingface.co
- RLHF: Reinforcement Learning from Human Feedbackhuyenchip.com
- Reward Hacking in Reinforcement Learning | Lil'Loglilianweng.github.io
- rlhfbook.com/book.pdfrlhfbook.com
- Rethinking the Role of PPO in RLHF – The Berkeley Artificial Intelligence Research Blogbair.berkeley.edu
- [2305.18290] Direct Preference Optimization: Your Language Model is Secretly a Reward Modelarxiv.org
- Energy-Based Preference Model Offers Better Offline Alignment than the Bradley-Terry Preference Modelarxiv.org