On the Algorithmic Bias of Aligning Large Language Models with RLHF: Preference Collapse and Matching Regularization
This is experimental HTML to improve accessibility. We invite you to report rendering errors. Use Alt+Y to toggle on accessible reporting links and Alt+Shift+Y to toggle off. Learn more about this project and help improve conversions. Accurately aligning large language models (LLMs) with human preferences is crucial for informing fair, economically sound, and statistically efficient decision-making processes. However, we argue that reinforcement learning from human feedback (RLHF)—the predominant approach for aligning LLMs with human preferences through a reward model—suffers from an inherent algorithmic bias due to its Kullback–Leibler-based regularization in optimization. In extreme cases, this bias could lead to a phenomenon we term preference collapse, where minority preferences are virtually disregarded. To mitigate this algorithmic bias, we introduce preference matching (PM) RLHF, a novel approach that provably aligns LLMs with the preference distribution of the reward model und
On the Algorithmic Bias of Aligning Large Language Models with RLHF: Preference Collapse and Matching Regularization Jiancong Xiao University of Pennsylvania. Emails: {jcxiao,qlong}@upenn.edu, suw@wharton.upenn.edu . Ziniu Li The Chinese University of Hong Kong, Shenzhen. Xingyu Xie National University of Singapore. Emily Getzen ∗ Cong Fang Peking University. Qi Long ∗ Weijie J. Su ∗ (May 26, 2024) Abstract Accurately aligning large language models (LLMs) with human preferences is crucial for informing fair, economically sound, and statistically efficient decision-making processes. However, we
Explore this link on the map →related reading
- [2305.18290] Direct Preference Optimization: Your Language Model is Secretly a Reward Modelarxiv.org
- RLHF & Post-Training Course by Nathan Lambertrlhfbook.com
- Nash Learning from Human Feedbackarxiv.org
- Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts | RLHFlowrlhflow.github.io
- RLHF: Reinforcement Learning from Human Feedbackhuyenchip.com
- State of RL for reasoning LLMs | A. Weersaweers.de
- rlhfbook.com/book.pdfrlhfbook.com
- [2503.10990] Statistical Impossibility and Possibility of Aligning LLMs with Human Preferences: From Condorcet Paradox to Nash Equilibriumarxiv.org
- [2402.14740] Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMsarxiv.org
- [2305.18290] Direct Preference Optimization: Your Language Model is Secretly a Reward Modelarxiv.org
- A vision researcher’s guide to some RL stuff: PPO & GRPO - Yuge (Jimmy) Shiyugeten.github.io
- Rethinking the Role of PPO in RLHF – The Berkeley Artificial Intelligence Research Blogbair.berkeley.edu