Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences?
This is experimental HTML to improve accessibility. We invite you to report rendering errors. Use Alt+Y to toggle on accessible reporting links and Alt+Shift+Y to toggle off. Learn more about this project and help improve conversions.
Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences? Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences? Paul Gölz Cornell University paulgoelz@cornell.edu Nika Haghtalab UC Berkeley nika@berkeley.edu Kunhe Yang UC Berkeley kunheyang@berkeley.edu Abstract After pre-training, large language models are aligned with human preferences based on pairwise comparisons. State-of-the-art alignment methods (such as PPO-based RLHF and DPO) are built on the assumption of aligning with a single preference model, despite being deployed in settings
Explore this link on the map →related reading
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- [2305.18290] Direct Preference Optimization: Your Language Model is Secretly a Reward Modelarxiv.org
- [2503.10990] Statistical Impossibility and Possibility of Aligning LLMs with Human Preferences: From Condorcet Paradox to Nash Equilibriumarxiv.org
- Breaking the Metric Voting Distortion Barrierarxiv.org
- Metric Distortion Under Probabilistic Votingarxiv.org
- On the Algorithmic Bias of Aligning Large Language Models with RLHF: Preference Collapse and Matching Regularizationarxiv.org
- Energy-Based Preference Model Offers Better Offline Alignment than the Bradley-Terry Preference Modelarxiv.org
- [2605.11134] Spurious Correlation Learning in Preference Optimization: Mechanisms, Consequences, and Mitigation via Tie Trainingarxiv.org
- Arrow's impossibility theorem - Wikipediaen.wikipedia.org
- Nash Learning from Human Feedbackarxiv.org
- [2402.01306] KTO: Model Alignment as Prospect Theoretic Optimizationarxiv.org
- Understanding the Performance Gap in Preference Learning: A Dichotomy of RLHF and DPOarxiv.org