✳flâneur — a map of the web's best reading
rlhfbook.com/book.pdf
rlhfbook.com · 93,316 words · saved by 1 readers
N/A
# link_1juwce72yay.pdf ## Metadata - PDFFormatVersion=1.7 - Language=en-US - IsLinearized=false - IsAcroFormPresent=false - IsXFAPresent=false - IsCollectionPresent=false - IsSignaturesPresent=false - Producer=pdfTeX-1.40.29 - Author=Nathan Lambert - Title=Reinforcement Learning from Human Feedback - Creator=LaTeX via pandoc - Keywords=RLHF, post-training, language models, LLMs, reward models, preference tuning, direct preference optimization, DPO, RLVR, reinforcement learning, AI alignment - CreationDate=D:20260623154302Z - ModDate=D:20260623154302Z - Trapped=False - Custom.PTEX.Fullbanner=Th
Explore this link on the map →saved by
related reading
- The State of Reinforcement Learning for LLM Reasoningmagazine.sebastianraschka.com
- A Guide to Reinforcement Learning Post-Training for LLMs: PPO, DPO, GRPO, and Beyondhuggingface.co
- RLHF: Reinforcement Learning from Human Feedbackhuyenchip.com
- RLHF & Post-Training Course by Nathan Lambertrlhfbook.com
- A vision researcher’s guide to some RL stuff: PPO & GRPO - Yuge (Jimmy) Shiyugeten.github.io
- RLHF Bookrlhfbook.com
- [2305.18290] Direct Preference Optimization: Your Language Model is Secretly a Reward Modelarxiv.org
- How RLHF actually works - by Nathan Lambertinterconnects.ai
- LLM Training: RLHF and Its Alternativesmagazine.sebastianraschka.com
- Reinforcement learning from human feedback - Wikipediaen.wikipedia.org
- [2402.14740] Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMsarxiv.org
- State of RL for reasoning LLMs | A. Weersaweers.de