flâneur — a map of the web's best reading

Rethinking the Role of PPO in RLHF – The Berkeley Artificial Intelligence Research Blog

bair.berkeley.edu · 1,695 words · saved by 1 readers

The BAIR Blog

Rethinking the Role of PPO in RLHF TL;DR : In RLHF, there’s tension between the reward learning phase, which uses human preference in the form of comparisons, and the RL fine-tuning phase, which optimizes a single, non-comparative reward. What if we performed RL in a comparative way? Figure 1: This diagram illustrates the difference between reinforcement learning from absolute feedback and relative feedback. By incorporating a new component - pairwise policy gradient, we can unify the reward modeling stage and RL stage, enabling direct updates based on pairwise responses. Large Language Models

Explore this link on the map →

related reading