How to Post-Train · RL Fundamentals Mini-Series
A no-BS series on what actually goes wrong in RL post-training - trajectory eyeballing, rubrics and verifiers, task design, environment quality - and how to fix it.
A running series of tutorials, opinionated pieces, and guides around what actually goes wrong in post-training. Everything from trajectory eyeballing, rubrics and verifiers, task design, to RL environment quality. Written from years in the trenches :). By Auriel · 5 posts RL Fundamentals Mini-Series 01 How to Eye Ball Trajectories: You’ve Never Spent Real Time with Your Model and We Can ALL Tell A no-BS guide for startups post-training their own models Live 02 RL Environment Harness Quality: Stop Shipping Low-Quality Harnesses and Calling It an “Environment” Flaky harnesses quietly…
saved by
related reading
- RL Pet Peeves Part 1 · Aurielaurielws.github.io
- Good QC for RL Dataseancai.com
- RLHF & Post-Training Course by Nathan Lambertrlhfbook.com
- Debugging Reinforcement Learning Systemsandyljones.com
- PostTrainBenchposttrainbench.com
- A Taxonomy of RL Environments for LLM Agentsleehanchung.github.io
- [2602.05910] Chunky Post-Training: Data Driven Failures of Generalizationarxiv.org
- aiesi_post-training_public.pdfkawine.github.io
- We need 3rd party Training-Run Assessments — LessWronglesswrong.com
- The upcoming GPT-3 moment for RL | Mechanize, Inc.mechanize.work
- Reinforcement Learning With Verifiable Rewards: How Data and Verifiers Shape RLVRsnorkel.ai
- RL Environments and RL for Science: Data Foundries and Multi-Agent Architecturesnewsletter.semianalysis.com