RLHF Book by Nathan Lambert
rlhfbook.com · 671 words · saved by 14 readers
A free online book and course on RLHF, preference tuning, reward models, RLVR, and post-training language models.
Web Version vs. Physical Book (Errata Fixes) The book will be re-printed roughly 2 and 6 months after the initial print in July 2026. This section tracks the differences between the web version and the physical book, and will be updated to note which improvements or fixes make it into which print version. Content additions and fixes: Expanded the canonical training recipes with MOPD and agentic post-training examples, plus new pipeline figures (Chapter 3) — #529. Added a short subsection on agentic evaluation (Chapter 16) — #492. Clarified the history of outcome reward models, fixed…
saved by
- Falah Rajput
- Rajan Agarwal
- Emma Guo
- Scott Langille
- Benedict Neo
- Trevor Trinh
- Rikard Saqe
- surya
- Akira Yoshiyama
- Neil Rathi
- John Zhang
- Ishaan Panigrahi
related reading
- RLHF & Post-Training Course by Nathan Lambertrlhfbook.com
- rlhfbook.com/book.pdfrlhfbook.com
- Introduction | RLHF and Post-Training Book by Nathan Lambertrlhfbook.com
- RLHF | John Lambertjohnwlambert.github.io
- RLHF: Reinforcement Learning from Human Feedbackhuyenchip.com
- State of RL for reasoning LLMs | A. Weersaweers.de
- RLHF Bookrlhfbook.com
- [2305.18290] Direct Preference Optimization: Your Language Model is Secretly a Reward Modelarxiv.org
- Reward Hacking in Reinforcement Learning | Lil'Loglilianweng.github.io
- GitHub - ash80/RLHF_in_notebooks: RLHF (Supervised fine-tuning, reward model, and PPO) step-by-step in 3 Jupyter notebooksgithub.com
- Fine-tune Llama 2 with DPOhuggingface.co
- Cornell RL Research Seminarxikronz.github.io