[2309.00267] RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback
Abstract:Reinforcement learning from human feedback (RLHF) has proven effective in aligning large language models (LLMs) with human preferences, but gathering high-quality preference labels is expensive. RL from AI Feedback (RLAIF), introduced in Bai et al., offers a promising alternative that trains the reward model (RM) on preferences generated by an off-the-shelf LLM. Across the tasks of summarization, helpful dialogue generation, and harmless dialogue generation, we show that RLAIF achieves comparable performance to RLHF. Furthermore, we take a step towards "self-improvement" by demonstrating that RLAIF can outperform a supervised fine-tuned baseline even when the AI labeler is the same size as the policy, or even the exact same checkpoint as the initial policy. Finally, we introduce direct-RLAIF (d-RLAIF) - a technique that circumvents RM training by obtaining rewards directly from an off-the-shelf LLM during RL, which achieves superior performance to canonical RLAIF. Our results suggest that RLAIF can achieve performance on-par with using human feedback, offering a potential solution to the scalability limitations of RLHF.
[2309.00267] RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback --> Computer Science > Computation and Language arXiv:2309.00267 (cs) [Submitted on 1 Sep 2023 ( v1 ), last revised 3 Sep 2024 (this version, v3)] Title: RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback Authors: Harrison Lee , Samrat Phatale , Hassan Mansoor , Thomas Mesnard , Johan Ferret , Kellie Lu , Colton Bishop , Ethan Hall , Victor Carbune , Abhinav Rastogi , Sushant Prakash View a PDF of the paper titled RLAIF vs. RLHF: Scaling Reinforcement Learning from
Explore this link on the map →saved by
related reading
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- RLHF: Reinforcement Learning from Human Feedbackhuyenchip.com
- RLHF & Post-Training Course by Nathan Lambertrlhfbook.com
- [2402.14740] Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMsarxiv.org
- Reinforcement learning from human feedback - Wikipediaen.wikipedia.org
- rlhfbook.com/book.pdfrlhfbook.com
- LLM Training: RLHF and Its Alternativesmagazine.sebastianraschka.com
- [2305.18290] Direct Preference Optimization: Your Language Model is Secretly a Reward Modelarxiv.org
- Thoughts on the impact of RLHF research — LessWronglesswrong.com
- How RLHF actually works - by Nathan Lambertinterconnects.ai
- State of RL for reasoning LLMs | A. Weersaweers.de
- Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts | RLHFlowrlhflow.github.io