✳flâneur — a map of the web's best reading
RLHF Book by Nathan Lambert
rlhfbook.com · 424 words · saved by 8 readers
A free online book and course on RLHF, preference tuning, reward models, RLVR, and post-training language models.
Explore this link on the map →