flâneur — a map of the web's best reading

Thoughts on the impact of RLHF research — LessWrong

lesswrong.com · 14,927 words · saved by 1 readers

In this post I’m going to describe my basic justification for working on RLHF in 2017-2020, which I still stand behind. I’ll discuss various argument…

x Thoughts on the impact of RLHF research — LessWrong RLHF AI Frontpage 255 Thoughts on the impact of RLHF research by paulfchristiano 25th Jan 2023 AI Alignment Forum 11 min read 102 255 Ω 111 In this post I’m going to describe my basic justification for working on RLHF in 2017-2020, which I still stand behind. I’ll discuss various arguments that RLHF research had an overall negative impact and explain why I don’t find them persuasive. I'll also clarify that I don't think research on RLHF is automatically net positive; alignment research should address real alignment problems, and we should r

Explore this link on the map →

saved by

related reading