flâneur

Introduction | RLHF and Post-Training Book by Nathan Lambert

rlhfbook.com · 5,952 words · saved by 1 readers

A first-principles introduction to RLHF, what it changes in language models, and how it became part of modern post-training.

Reinforcement learning from human feedback (RLHF) is a technique used to incorporate human information into AI systems. RLHF emerged primarily as a method to solve hard-to-specify problems. With systems that are designed to be used by humans directly, such problems emerge all the time due to the often inexpressible nature of an individual’s preferences. This encompasses every domain of content and interaction with a digital system. RLHF’s early applications were often in control problems and other traditional domains for reinforcement learning (RL), where the goal is to optimize a specific…

saved by

related reading