[2402.14740] Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs
Abstract:AI alignment in the shape of Reinforcement Learning from Human Feedback (RLHF) is increasingly treated as a crucial ingredient for high performance large language models. Proximal Policy Optimization (PPO) has been positioned by recent literature as the canonical method for the RL part of RLHF. However, it involves both high computational cost and sensitive hyperparameter tuning. We posit that most of the motivational principles that led to the development of PPO are less of a practical concern in RLHF and advocate for a less computationally expensive method that preserves and even increases performance. We revisit the formulation of alignment from human preferences in the context of RL. Keeping simplicity as a guiding principle, we show that many components of PPO are unnecessary in an RLHF context and that far simpler REINFORCE-style optimization variants outperform both PPO and newly proposed "RL-free" methods such as DPO and RAFT. Our work suggests that careful adaptation to LLMs alignment characteristics enables benefiting from online RL optimization at low cost.
[2402.14740] Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs --> Computer Science > Machine Learning arXiv:2402.14740 (cs) [Submitted on 22 Feb 2024 ( v1 ), last revised 26 Feb 2024 (this version, v2)] Title: Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs Authors: Arash Ahmadian , Chris Cremer , Matthias Gallé , Marzieh Fadaee , Julia Kreutzer , Olivier Pietquin , Ahmet Üstün , Sara Hooker View a PDF of the paper titled Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Hum
Explore this link on the map →saved by
related reading
- Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts | RLHFlowrlhflow.github.io
- RLHF & Post-Training Course by Nathan Lambertrlhfbook.com
- RLHF: Reinforcement Learning from Human Feedbackhuyenchip.com
- State of RL for reasoning LLMs | A. Weersaweers.de
- [2305.18290] Direct Preference Optimization: Your Language Model is Secretly a Reward Modelarxiv.org
- rlhfbook.com/book.pdfrlhfbook.com
- A vision researcher’s guide to some RL stuff: PPO & GRPO - Yuge (Jimmy) Shiyugeten.github.io
- Reinforcement learning from human feedback - Wikipediaen.wikipedia.org
- [2309.00267] RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedbackarxiv.org
- LLM Training: RLHF and Its Alternativesmagazine.sebastianraschka.com
- Rethinking the Role of PPO in RLHF – The Berkeley Artificial Intelligence Research Blogbair.berkeley.edu
- The N Implementation Details of RLHF with PPO | ICLR Blogposts 2024iclr-blogposts.github.io