flâneur

How to Post-Train · RL Fundamentals Mini-Series

howtoposttrain.com · 322 words · saved by 1 readers

A no-BS series on what actually goes wrong in RL post-training - trajectory eyeballing, rubrics and verifiers, task design, environment quality - and how to fix it.

A running series of tutorials, opinionated pieces, and guides around what actually goes wrong in post-training. Everything from trajectory eyeballing, rubrics and verifiers, task design, to RL environment quality. Written from years in the trenches :). By Auriel · 5 posts RL Fundamentals Mini-Series 01 How to Eye Ball Trajectories: You’ve Never Spent Real Time with Your Model and We Can ALL Tell A no-BS guide for startups post-training their own models Live 02 RL Environment Harness Quality: Stop Shipping Low-Quality Harnesses and Calling It an “Environment” Flaky harnesses quietly…

saved by

related reading