flâneur

RLHF Book by Nathan Lambert

rlhfbook.com · 671 words · saved by 14 readers

A free online book and course on RLHF, preference tuning, reward models, RLVR, and post-training language models.

Web Version vs. Physical Book (Errata Fixes) The book will be re-printed roughly 2 and 6 months after the initial print in July 2026. This section tracks the differences between the web version and the physical book, and will be updated to note which improvements or fixes make it into which print version. Content additions and fixes: Expanded the canonical training recipes with MOPD and agentic post-training examples, plus new pipeline figures (Chapter 3) — #529. Added a short subsection on agentic evaluation (Chapter 16) — #492. Clarified the history of outcome reward models, fixed…

saved by

related reading