flâneur

RLHF | John Lambert

johnwlambert.github.io · 3,835 words · saved by 2 readers

Reinforcement Learning from Human Feedback

Table of Contents: Overview Reward Model Training PPO Direct Preference Optimization (DPO) DPO: Deriving the Optimum of the KL-Constrained Reward Maximization Objective DPO: Solving for Reward, Using the Optimal Policy DPO: Revealing the DPO Loss Function DPO: Deriving the DPO Objective Under the Bradley-Terry Model GRPO Post-Training in the Post-RLHF Era Overview Reinforcement learning from human feedback (RLHF) align models (and language modeling objective) with users’ values to be truthful, non-toxic, and helpful to the user, through the use of a trained reward model. Because…

saved by

related reading