flâneur — a map of the web's best reading

Preference Data | RLHF Book by Nathan Lambert

rlhfbook.com · 4,733 words · saved by 1 readers

The Reinforcement Learning from Human Feedback Book

--> RLHF Book --> Preference Data | RLHF and Post-Training Book by Nathan Lambert Reinforcement Learning from Human Feedback A short introduction to RLHF and post-training focused on language models. Nathan Lambert Lecture 8: On “Preferences” and Preference Data Preference Data Preference data is the engine of preference fine-tuning and reinforcement learning from human feedback. The core problem we’ve been trying to solve with RLHF is that we cannot precisely model human rewards and preferences for AI models’ outputs – that is, write clearly defined loss functions to optimize against – so pre

Explore this link on the map →

related reading