Thoughts on the impact of RLHF research - AI Alignment Forum
In this post I’m going to describe my basic justification for working on RLHF in 2017-2020, which I still stand behind. I’ll discuss various arguments that RLHF research had an overall negative impac…
x Thoughts on the impact of RLHF research — AI Alignment Forum RLHF AI Frontpage 111 Thoughts on the impact of RLHF research by paulfchristiano 25th Jan 2023 11 min read 102 111 In this post I’m going to describe my basic justification for working on RLHF in 2017-2020, which I still stand behind. I’ll discuss various arguments that RLHF research had an overall negative impact and explain why I don’t find them persuasive. I'll also clarify that I don't think research on RLHF is automatically net positive; alignment research should address real alignment problems, and we should reject a vague as
saved by
related reading
- Thoughts on the impact of RLHF research — LessWronglesswrong.com
- RLHF: Reinforcement Learning from Human Feedbackhuyenchip.com
- AI in 2025: gestalt — LessWronglesswrong.com
- Alignment is not solved but it increasingly looks solvablealigned.substack.com
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- rlhfbook.com/book.pdfrlhfbook.com
- AI #23: Fundamental Problems with RLHFthezvi.substack.com
- Introduction | RLHF and Post-Training Book by Nathan Lambertrlhfbook.com
- The N Implementation Details of RLHF with PPO | ICLR Blogposts 2024iclr-blogposts.github.io
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- Why I’m optimistic about our alignment approachaligned.substack.com
- How RLHF actually works - by Nathan Lambertinterconnects.ai