flâneur — a map of the web's best reading

RLHF: Reinforcement Learning from Human Feedback

huyenchip.com · 4,729 words · saved by 9 readers

In literature discussing why ChatGPT is able to capture so much of our imagination, I often come across two narratives:

[ LinkedIn discussion , Twitter thread ] In literature discussing why ChatGPT is able to capture so much of our imagination, I often come across two narratives: Scale: throwing more data and compute at it. UX: moving from a prompt interface to a more natural chat interface. One narrative that is often glossed over is the incredible technical creativity that went into making models like ChatGPT work. One such cool idea is RLHF (Reinforcement Learning from Human Feedback): incorporating reinforcement learning and human feedback into NLP. RL has been notoriously difficult to work with, and theref

Explore this link on the map →

saved by

related reading