flâneur — a map of the web's best reading

How RLHF actually works - by Nathan Lambert - Interconnects

interconnects.ai · 2,526 words · saved by 1 readers

Why RLHF may still win out and why we haven't seen it yet in open-source.

How RLHF actually works The proven formula for RLHF and when we will see it in open-source. Nathan Lambert Jun 21, 2023 55 3 2 Share The question I still get the most is "Why does reinforcement learning from human feedback (RLHF) work?" Until last week, my answer was still "no one knows." We are starting to get some answers. RLHF ultimately will work in the long term (with language models and elsewhere) when two conditions are met. First, there needs to be some signal that applying vanilla supervised learning only does not work — in this case the pairwise preference data. Second, the less impo

Explore this link on the map →

related reading