Thoughts on the impact of RLHF research — LessWrong
In this post I’m going to describe my basic justification for working on RLHF in 2017-2020, which I still stand behind. I’ll discuss various argument…
x Thoughts on the impact of RLHF research — LessWrong RLHF AI Frontpage 255 Thoughts on the impact of RLHF research by paulfchristiano 25th Jan 2023 AI Alignment Forum 11 min read 102 255 Ω 111 In this post I’m going to describe my basic justification for working on RLHF in 2017-2020, which I still stand behind. I’ll discuss various arguments that RLHF research had an overall negative impact and explain why I don’t find them persuasive. I'll also clarify that I don't think research on RLHF is automatically net positive; alignment research should address real alignment problems, and we should r
saved by
related reading
- Thoughts on the impact of RLHF research — AI Alignment Forumalignmentforum.org
- RLHF: Reinforcement Learning from Human Feedbackhuyenchip.com
- rlhfbook.com/book.pdfrlhfbook.com
- Introduction | RLHF and Post-Training Book by Nathan Lambertrlhfbook.com
- RLHF & Post-Training Course by Nathan Lambertrlhfbook.com
- AI in 2025: gestalt — LessWronglesswrong.com
- How can LLM RL Work Despite Information-Theoretic Inefficiencyberen.io
- AI #23: Fundamental Problems with RLHFthezvi.substack.com
- The N Implementation Details of RLHF with PPO | ICLR Blogposts 2024iclr-blogposts.github.io
- Open Problems of Reinforcement Learningarxiv.org
- How RLHF actually works - by Nathan Lambertinterconnects.ai
- RLHF | John Lambertjohnwlambert.github.io