Nash Learning from Human Feedback
This is experimental HTML to improve accessibility. We invite you to report rendering errors. Use Alt+Y to toggle on accessible reporting links and Alt+Shift+Y to toggle off. Learn more about this project and help improve conversions. Reinforcement learning from human feedback (RLHF) has emerged as the main paradigm for aligning large language models (LLMs) with human preferences. Traditionally, RLHF involves the initial step of learning a reward model from pairwise human feedback, i.e., expressed as preferences between pairs of text generations. Subsequently, the LLM’s policy is fine-tuned to maximize the reward through a reinforcement learning algorithm. In this study, we introduce an alternative pipeline for the fine-tuning of LLMs using pairwise human feedback. Our approach entails the initial learning of a pairwise preference model, which is conditioned on two inputs (instead of a single input in the case of a reward model) given a prompt, followed by the pursuit of a policy that
Nash Learning from Human Feedback Rémi Munos Michal Valko Daniele Calandriello Mohammad Gheshlaghi Azar Mark Rowland Daniel Guo Yunhao Tang Matthieu Geist Thomas Mesnard Côme Fiegel Andrea Michi Marco Selvi Sertan Girgin Nikola Momchev Olivier Bachem Daniel J. Mankowitz Doina Precup Bilal Piot Abstract Reinforcement learning from human feedback (RLHF) has emerged as the main paradigm for aligning large language models (LLMs) with human preferences. Traditionally, RLHF involves the initial step of learning a reward model from pairwise human feedback, i.e., expressed as preferences between pairs
Explore this link on the map →related reading
- [2305.18290] Direct Preference Optimization: Your Language Model is Secretly a Reward Modelarxiv.org
- RLHF & Post-Training Course by Nathan Lambertrlhfbook.com
- RLHF: Reinforcement Learning from Human Feedbackhuyenchip.com
- Reinforcement learning from human feedback - Wikipediaen.wikipedia.org
- [2503.10990] Statistical Impossibility and Possibility of Aligning LLMs with Human Preferences: From Condorcet Paradox to Nash Equilibriumarxiv.org
- On the Algorithmic Bias of Aligning Large Language Models with RLHF: Preference Collapse and Matching Regularizationarxiv.org
- rlhfbook.com/book.pdfrlhfbook.com
- RLHF Bookrlhfbook.com
- Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts | RLHFlowrlhflow.github.io
- [2305.18290] Direct Preference Optimization: Your Language Model is Secretly a Reward Modelarxiv.org
- RLHF Bookrlhfbook.com
- [2402.14740] Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMsarxiv.org