Nash Learning from Human Feedback
This is experimental HTML to improve accessibility. We invite you to report rendering errors. Use Alt+Y to toggle on accessible reporting links and Alt+Shift+Y to toggle off. Learn more about this project and help improve conversions. Reinforcement learning from human feedback (RLHF) has emerged as the main paradigm for aligning large language models (LLMs) with human preferences. Traditionally, RLHF involves the initial step of learning a reward model from pairwise human feedback, i.e., expressed as preferences between pairs of text generations. Subsequently, the LLM’s policy is fine-tuned to maximize the reward through a reinforcement learning algorithm. In this study, we introduce an alternative pipeline for the fine-tuning of LLMs using pairwise human feedback. Our approach entails the initial learning of a pairwise preference model, which is conditioned on two inputs (instead of a single input in the case of a reward model) given a prompt, followed by the pursuit of a policy that
Nash Learning from Human Feedback Rémi Munos Michal Valko Daniele Calandriello Mohammad Gheshlaghi Azar Mark Rowland Daniel Guo Yunhao Tang Matthieu Geist Thomas Mesnard Côme Fiegel Andrea Michi Marco Selvi Sertan Girgin Nikola Momchev Olivier Bachem Daniel J. Mankowitz Doina Precup Bilal Piot Abstract Reinforcement learning from human feedback (RLHF) has emerged as the main paradigm for aligning large language models (LLMs) with human preferences. Traditionally, RLHF involves the initial step of learning a reward model from pairwise human feedback, i.e., expressed as preferences between pairs
related reading
- [2305.18290] Direct Preference Optimization: Your Language Model is Secretly a Reward Modelarxiv.org
- RLHF & Post-Training Course by Nathan Lambertrlhfbook.com
- RLHF | John Lambertjohnwlambert.github.io
- RLHF: Reinforcement Learning from Human Feedbackhuyenchip.com
- Reinforcement learning from human feedback - Wikipediaen.wikipedia.org
- Direct Nash Optimization: Teaching Language Models to Self-Improve with General Preferencesarxiv.org
- [2503.10990] Statistical Impossibility and Possibility of Aligning LLMs with Human Preferences: From Condorcet Paradox to Nash Equilibriumarxiv.org
- On the Algorithmic Bias of Aligning Large Language Models with RLHF: Preference Collapse and Matching Regularizationarxiv.org
- Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts | RLHFlowrlhflow.github.io
- rlhfbook.com/book.pdfrlhfbook.com
- RLHF Bookrlhfbook.com
- Introduction | RLHF and Post-Training Book by Nathan Lambertrlhfbook.com