flâneur — a map of the web's best reading

Reinforcement Learning | RLHF and Post-Training Book by Nathan Lambert

rlhfbook.com · 15,424 words · saved by 1 readers

Policy gradient methods for RLHF and LLM post-training, including PPO, REINFORCE, RLOO, GRPO, and implementation details.

--> RLHF Book --> Reinforcement Learning | RLHF and Post-Training Book by Nathan Lambert Reinforcement Learning from Human Feedback A short introduction to RLHF and post-training focused on language models. Nathan Lambert Lecture 3: Understanding Policy Gradient Algorithms for RL on LLMs Lecture 4: Implementing RL Algorithms for LLMs Reinforcement Learning In the RLHF process, the reinforcement learning algorithm slowly updates the model’s weights with respect to feedback from a reward model. The policy – the model being trained – generates completions to prompts in the training set, then the

Explore this link on the map →

saved by

related reading