[2503.10990] Statistical Impossibility and Possibility of Aligning LLMs with Human Preferences: From Condorcet Paradox to Nash Equilibrium
Abstract:Aligning large language models (LLMs) with diverse human preferences is critical for ensuring fairness and informed outcomes when deploying these models for decision-making. In this paper, we seek to uncover fundamental statistical limits concerning aligning LLMs with human preferences, with a focus on the probabilistic representation of human preferences and the preservation of diverse preferences in aligned LLMs. We first show that human preferences can be represented by a reward model if and only if the preference among LLM-generated responses is free of any Condorcet cycle. Moreover, we prove that Condorcet cycles exist with probability converging to one exponentially fast under a general probabilistic preference model called the Luce model, thereby demonstrating the impossibility of fully aligning human preferences using reward-based approaches such as reinforcement learning from human feedback. Next, we explore the conditions under which LLMs would employ mixed strategies -- meaning they do not collapse to a single response -- when aligned in the limit using a non-reward-based approach, such as Nash learning from human feedback. We identify a necessary and sufficient condition for mixed strategies: the absence of a response that is preferred over all others by a majority. As a blessing, we prove that this condition holds with high probability under the Luce model, thereby highlighting the statistical possibility of preserving minority preferences without explicit regularization in aligning LLMs.
# link_1sofd6x6er1.pdf ## Metadata - PDFFormatVersion=1.7 - IsLinearized=false - IsAcroFormPresent=false - IsXFAPresent=false - IsCollectionPresent=false - IsSignaturesPresent=false - Author=Kaizhao Liu; Qi Long; Zhekun Shi; Weijie J. Su; Jiancong Xiao - Creator=arXiv GenPDF (tex2pdf:a6404ea) - Custom.DOI=https://doi.org/10.48550/arXiv.2503.10990 - Custom.License=http://arxiv.org/licenses/nonexclusive-distrib/1.0/ - Custom.PTEX.Fullbanner=This is pdfTeX, Version 3.141592653-2.6-1.40.28 (TeX Live 2025) kpathsea version 6.4.1 - Custom.arXivID=https://arxiv.org/abs/2503.10990v2 - Producer=pikepdf
saved by
related reading
- Nash Learning from Human Feedbackarxiv.org
- [2305.18290] Direct Preference Optimization: Your Language Model is Secretly a Reward Modelarxiv.org
- Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts | RLHFlowrlhflow.github.io
- On the Algorithmic Bias of Aligning Large Language Models with RLHF: Preference Collapse and Matching Regularizationarxiv.org
- [2506.00751] Alignment Revisited: Are Large Language Models Consistent in Stated and Revealed Preferences?arxiv.org
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- 2307.12950.pdfarxiv.org
- [2302.08582] Pretraining Language Models with Human Preferencesarxiv.org
- RLHF: Reinforcement Learning from Human Feedbackhuyenchip.com
- RLHF Bookrlhfbook.com
- Direct Nash Optimization: Teaching Language Models to Self-Improve with General Preferencesarxiv.org
- The two types of LLM preferencesnewsletter.danielpaleka.com