flâneur — a map of the web's best reading

Thoughts on the impact of RLHF research - AI Alignment Forum

alignmentforum.org · 11,905 words · saved by 1 readers

In this post I’m going to describe my basic justification for working on RLHF in 2017-2020, which I still stand behind. I’ll discuss various arguments that RLHF research had an overall negative impac…

x Thoughts on the impact of RLHF research — AI Alignment Forum RLHF AI Frontpage 111 Thoughts on the impact of RLHF research by paulfchristiano 25th Jan 2023 11 min read 102 111 In this post I’m going to describe my basic justification for working on RLHF in 2017-2020, which I still stand behind. I’ll discuss various arguments that RLHF research had an overall negative impact and explain why I don’t find them persuasive. I'll also clarify that I don't think research on RLHF is automatically net positive; alignment research should address real alignment problems, and we should reject a vague as

Explore this link on the map →

saved by

related reading