Thoughts on the impact of RLHF research - AI Alignment Forum
In this post I’m going to describe my basic justification for working on RLHF in 2017-2020, which I still stand behind. I’ll discuss various arguments that RLHF research had an overall negative impac…
x Thoughts on the impact of RLHF research — AI Alignment Forum RLHF AI Frontpage 111 Thoughts on the impact of RLHF research by paulfchristiano 25th Jan 2023 11 min read 102 111 In this post I’m going to describe my basic justification for working on RLHF in 2017-2020, which I still stand behind. I’ll discuss various arguments that RLHF research had an overall negative impact and explain why I don’t find them persuasive. I'll also clarify that I don't think research on RLHF is automatically net positive; alignment research should address real alignment problems, and we should reject a vague as
Explore this link on the map →saved by
related reading
- Thoughts on the impact of RLHF research — LessWronglesswrong.com
- RLHF: Reinforcement Learning from Human Feedbackhuyenchip.com
- AI in 2025: gestalt — LessWronglesswrong.com
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- rlhfbook.com/book.pdfrlhfbook.com
- The N Implementation Details of RLHF with PPO | ICLR Blogposts 2024iclr-blogposts.github.io
- Why I’m optimistic about our alignment approachaligned.substack.com
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- How RLHF actually works - by Nathan Lambertinterconnects.ai
- [2309.00267] RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedbackarxiv.org
- Proposal: Using Monte Carlo tree search instead of RLHF for alignment research — LessWronglesswrong.com
- I am worried about near-term non-LLM AI developments — LessWronglesswrong.com