rlhfbook.com/book.pdf
rlhfbook.com · 93,316 words · saved by 1 readers
N/A
# link_1juwce72yay.pdf ## Metadata - PDFFormatVersion=1.7 - Language=en-US - IsLinearized=false - IsAcroFormPresent=false - IsXFAPresent=false - IsCollectionPresent=false - IsSignaturesPresent=false - Producer=pdfTeX-1.40.29 - Author=Nathan Lambert - Title=Reinforcement Learning from Human Feedback - Creator=LaTeX via pandoc - Keywords=RLHF, post-training, language models, LLMs, reward models, preference tuning, direct preference optimization, DPO, RLVR, reinforcement learning, AI alignment - CreationDate=D:20260623154302Z - ModDate=D:20260623154302Z - Trapped=False - Custom.PTEX.Fullbanner=Th
saved by
related reading
- The State of Reinforcement Learning for LLM Reasoningmagazine.sebastianraschka.com
- A Guide to Reinforcement Learning Post-Training for LLMs: PPO, DPO, GRPO, and Beyondhuggingface.co
- Introduction | RLHF and Post-Training Book by Nathan Lambertrlhfbook.com
- RLHF: Reinforcement Learning from Human Feedbackhuyenchip.com
- RLHF & Post-Training Course by Nathan Lambertrlhfbook.com
- RLHF | John Lambertjohnwlambert.github.io
- RLHF Bookrlhfbook.com
- [2305.18290] Direct Preference Optimization: Your Language Model is Secretly a Reward Modelarxiv.org
- How RLHF actually works - by Nathan Lambertinterconnects.ai
- LLM Training: RLHF and Its Alternativesmagazine.sebastianraschka.com
- Reinforcement learning from human feedback - Wikipediaen.wikipedia.org
- Fine-tune Llama 2 with DPOhuggingface.co