flâneur — a map of the web's best reading

Rejection Sampling | RLHF Book by Nathan Lambert

rlhfbook.com · 2,454 words · saved by 1 readers

The Reinforcement Learning from Human Feedback Book

--> RLHF Book --> Rejection Sampling | RLHF and Post-Training Book by Nathan Lambert Reinforcement Learning from Human Feedback A short introduction to RLHF and post-training focused on language models. Nathan Lambert Lecture 2: IFT, Reward Modeling, Rejection Sampling (Chap. 4, 5, & 9) Rejection Sampling Rejection Sampling (RS) is one of the most widely used yet least documented methods in preference fine-tuning. Many prominent RLHF papers use it as a core component of their training pipeline, yet no canonical implementation or explanation of why it works so well exists. RS can be applied at

Explore this link on the map →

related reading