flâneur — a map of the web's best reading

[2309.00267] RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback

arxiv.org · 729 words · saved by 1 readers

Abstract:Reinforcement learning from human feedback (RLHF) has proven effective in aligning large language models (LLMs) with human preferences, but gathering high-quality preference labels is expensive. RL from AI Feedback (RLAIF), introduced in Bai et al., offers a promising alternative that trains the reward model (RM) on preferences generated by an off-the-shelf LLM. Across the tasks of summarization, helpful dialogue generation, and harmless dialogue generation, we show that RLAIF achieves comparable performance to RLHF. Furthermore, we take a step towards "self-improvement" by demonstrating that RLAIF can outperform a supervised fine-tuned baseline even when the AI labeler is the same size as the policy, or even the exact same checkpoint as the initial policy. Finally, we introduce direct-RLAIF (d-RLAIF) - a technique that circumvents RM training by obtaining rewards directly from an off-the-shelf LLM during RL, which achieves superior performance to canonical RLAIF. Our results suggest that RLAIF can achieve performance on-par with using human feedback, offering a potential solution to the scalability limitations of RLHF.

[2309.00267] RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback --> Computer Science > Computation and Language arXiv:2309.00267 (cs) [Submitted on 1 Sep 2023 ( v1 ), last revised 3 Sep 2024 (this version, v3)] Title: RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback Authors: Harrison Lee , Samrat Phatale , Hassan Mansoor , Thomas Mesnard , Johan Ferret , Kellie Lu , Colton Bishop , Ethan Hall , Victor Carbune , Abhinav Rastogi , Sushant Prakash View a PDF of the paper titled RLAIF vs. RLHF: Scaling Reinforcement Learning from

Explore this link on the map →

saved by

related reading