Energy-Based Preference Model Offers Better Offline Alignment than the Bradley-Terry Preference Model
This is experimental HTML to improve accessibility. We invite you to report rendering errors. Use Alt+Y to toggle on accessible reporting links and Alt+Shift+Y to toggle off. Learn more about this project and help improve conversions.
Energy-Based Preference Model Offers Better Offline Alignment than the Bradley-Terry Preference Model Yuzhong Hong Hanshan Zhang Junwei Bao Hongfei Jiang Yang Song Abstract Since the debut of DPO, it has been shown that aligning a target LLM with human preferences via the KL-constrained RLHF loss is mathematically equivalent to a special kind of reward modeling task. Concretely, the task requires: 1) using the target LLM to parameterize the reward model, and 2) tuning the reward model so that it has a 1:1 linear relationship with the true reward. However, we identify a significant issue: the D
Explore this link on the map →related reading
- [2305.18290] Direct Preference Optimization: Your Language Model is Secretly a Reward Modelarxiv.org
- Understanding the Performance Gap in Preference Learning: A Dichotomy of RLHF and DPOarxiv.org
- RLHF & Post-Training Course by Nathan Lambertrlhfbook.com
- On the Algorithmic Bias of Aligning Large Language Models with RLHF: Preference Collapse and Matching Regularizationarxiv.org
- State of RL for reasoning LLMs | A. Weersaweers.de
- Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts | RLHFlowrlhflow.github.io
- RLHF: Reinforcement Learning from Human Feedbackhuyenchip.com
- Prompt-to-Leaderboardarxiv.org
- A Guide to Reinforcement Learning Post-Training for LLMs: PPO, DPO, GRPO, and Beyondhuggingface.co
- Nash Learning from Human Feedbackarxiv.org
- RLHF Bookrlhfbook.com
- Is DPO Superior to PPO for LLM Alignment? A Comprehensive Studyarxiv.org