How RLHF actually works - by Nathan Lambert - Interconnects
interconnects.ai · 2,526 words · saved by 1 readers
Why RLHF may still win out and why we haven't seen it yet in open-source.
How RLHF actually works The proven formula for RLHF and when we will see it in open-source. Nathan Lambert Jun 21, 2023 55 3 2 Share The question I still get the most is "Why does reinforcement learning from human feedback (RLHF) work?" Until last week, my answer was still "no one knows." We are starting to get some answers. RLHF ultimately will work in the long term (with language models and elsewhere) when two conditions are met. First, there needs to be some signal that applying vanilla supervised learning only does not work — in this case the pairwise preference data. Second, the less impo
related reading
- RLHF: Reinforcement Learning from Human Feedbackhuyenchip.com
- rlhfbook.com/book.pdfrlhfbook.com
- RLHF & Post-Training Course by Nathan Lambertrlhfbook.com
- Introduction | RLHF and Post-Training Book by Nathan Lambertrlhfbook.com
- [2305.18290] Direct Preference Optimization: Your Language Model is Secretly a Reward Modelarxiv.org
- RLHF | John Lambertjohnwlambert.github.io
- RLHF Bookrlhfbook.com
- Thoughts on the impact of RLHF research — LessWronglesswrong.com
- [2309.00267] RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedbackarxiv.org
- Open Problems of Reinforcement Learningarxiv.org
- RLHF Bookrlhfbook.com
- Reinforcement learning from human feedback - Wikipediaen.wikipedia.org