Thoughts on the impact of RLHF research — LessWrong
In this post I’m going to describe my basic justification for working on RLHF in 2017-2020, which I still stand behind. I’ll discuss various argument…
x Thoughts on the impact of RLHF research — LessWrong RLHF AI Frontpage 255 Thoughts on the impact of RLHF research by paulfchristiano 25th Jan 2023 AI Alignment Forum 11 min read 102 255 Ω 111 In this post I’m going to describe my basic justification for working on RLHF in 2017-2020, which I still stand behind. I’ll discuss various arguments that RLHF research had an overall negative impact and explain why I don’t find them persuasive. I'll also clarify that I don't think research on RLHF is automatically net positive; alignment research should address real alignment problems, and we should r
Explore this link on the map →saved by
related reading
- Thoughts on the impact of RLHF research — AI Alignment Forumalignmentforum.org
- RLHF: Reinforcement Learning from Human Feedbackhuyenchip.com
- rlhfbook.com/book.pdfrlhfbook.com
- RLHF & Post-Training Course by Nathan Lambertrlhfbook.com
- AI in 2025: gestalt — LessWronglesswrong.com
- The N Implementation Details of RLHF with PPO | ICLR Blogposts 2024iclr-blogposts.github.io
- How RLHF actually works - by Nathan Lambertinterconnects.ai
- Reward Hacking in Reinforcement Learning | Lil'Loglilianweng.github.io
- State of RL for reasoning LLMs | A. Weersaweers.de
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- [2309.00267] RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedbackarxiv.org
- [2402.14740] Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMsarxiv.org