✳flâneur — a map of the web's best reading
RLHF & Post-Training Course by Nathan Lambert
rlhfbook.com · 1,233 words · saved by 14 readers
Free course lectures on RLHF, reward models, preference tuning, RLVR, and modern LLM post-training.
RLHF & Post-Training Course by Nathan Lambert Course A full course accompanying the book with added resources and other lectures I've given. The slide decks are usually built with Claude Opus, Colloquium , and substantial human revisions. Welcome to the Course Introduction and overview of what you'll learn Watch Prerequisites - what to know before starting. Primary Material - the core lecture series that follows the book chapter by chapter, with recordings, slides, PDFs, and source. Extra Resources - recommended books, external RL courses, and Nathan's own talks paired with the chapters they g
Explore this link on the map →saved by
- Winnie Xu
- Anson Yu
- Mathurah Ravigulan
- Gary Xin
- Linda
- Benedict Neo
- Dhruv Sheth
- Jonathan Xu
- Faisal Sayed
- Jackson Mowatt Gok
- Nicklaus Tran
- 7165
related reading
- RLHF: Reinforcement Learning from Human Feedbackhuyenchip.com
- rlhfbook.com/book.pdfrlhfbook.com
- State of RL for reasoning LLMs | A. Weersaweers.de
- [2305.18290] Direct Preference Optimization: Your Language Model is Secretly a Reward Modelarxiv.org
- [2402.14740] Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMsarxiv.org
- GenAI Handbookgenai-handbook.github.io
- A vision researcher’s guide to some RL stuff: PPO & GRPO - Yuge (Jimmy) Shiyugeten.github.io
- Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts | RLHFlowrlhflow.github.io
- RLHF Bookrlhfbook.com
- LLM Training: RLHF and Its Alternativesmagazine.sebastianraschka.com
- The N Implementation Details of RLHF with PPO | ICLR Blogposts 2024iclr-blogposts.github.io
- A Guide to Reinforcement Learning Post-Training for LLMs: PPO, DPO, GRPO, and Beyondhuggingface.co