flâneur — a map of the web's best reading

ash80/RLHF_in_notebooks: RLHF (Supervised fine-tuning, reward model, and PPO) step-by-step in 3 Jupyter notebooks ·

github.com · 594 words · saved by 1 readers

RLHF (Supervised fine-tuning, reward model, and PPO) step-by-step in 3 Jupyter notebooks

Explore this link on the map →

saved by

related reading