RLHF: Reinforcement Learning from Human Feedback
huyenchip.com · 4,729 words · saved by 9 readers
In literature discussing why ChatGPT is able to capture so much of our imagination, I often come across two narratives:
[ LinkedIn discussion , Twitter thread ] In literature discussing why ChatGPT is able to capture so much of our imagination, I often come across two narratives: Scale: throwing more data and compute at it. UX: moving from a prompt interface to a more natural chat interface. One narrative that is often glossed over is the incredible technical creativity that went into making models like ChatGPT work. One such cool idea is RLHF (Reinforcement Learning from Human Feedback): incorporating reinforcement learning and human feedback into NLP. RL has been notoriously difficult to work with, and theref
saved by
- Eric Chang
- Yung-Hsuan Yang
- Micah Carroll
- Soham Shah
- Evan Cater
- Karan Dalal
- Manh Duc Hoang
- Hiệp Nguyễn Tuấn
- Danish
related reading
- rlhfbook.com/book.pdfrlhfbook.com
- RLHF & Post-Training Course by Nathan Lambertrlhfbook.com
- Introduction | RLHF and Post-Training Book by Nathan Lambertrlhfbook.com
- Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts | RLHFlowrlhflow.github.io
- RLHF | John Lambertjohnwlambert.github.io
- GitHub - ash80/RLHF_in_notebooks: RLHF (Supervised fine-tuning, reward model, and PPO) step-by-step in 3 Jupyter notebooksgithub.com
- LLM Training: RLHF and Its Alternativesmagazine.sebastianraschka.com
- How can LLM RL Work Despite Information-Theoretic Inefficiencyberen.io
- [2305.18290] Direct Preference Optimization: Your Language Model is Secretly a Reward Modelarxiv.org
- [2309.00267] RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedbackarxiv.org
- State of RL for reasoning LLMs | A. Weersaweers.de
- 2401.10020.pdfarxiv.org