flâneur — a map of the web's best reading

RLHF Book by Nathan Lambert

rlhfbook.com · 424 words · saved by 8 readers

A free online book and course on RLHF, preference tuning, reward models, RLVR, and post-training language models.

Explore this link on the map →

saved by