RL Pet Peeves Part 1 · Auriel
aurielws.github.io · 3,448 words · saved by 2 readers
My personal do's and don'ts for startups post-training their own model. A guide to trajectory eyeballing, harness debugging, and RL failure modes.
RL Pet Peeves Part 1 · Auriel This mini-series is a collection of rants on RL from my POV. Pure, unfiltered, first-person opinions from someone who's spent years deep in the trenches of pre-training, post-training/fine-tuning, inference time, and every layer of the stack for models from small distilled models (Pixel Real Tone base model) to frontier systems (Gemini + Nano Banana + Human Detection Models that powered Google Search, Waymo, Vertex AI). I've eyeballed thousands of trajectories, judged parametric wins and losses until my eyes bled at 2AM, and sat through more "data" pitches
saved by
related reading
- How to Post-Train · RL Fundamentals Mini-Serieshowtoposttrain.com
- Demystifying evals for AI agents \ Anthropicanthropic.com
- Your AI Product Needs Evals – Hamel's Blog - Hamel Husainhamel.dev
- Foundation Models for Oversight | Transluce AItransluce.org
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- Agent Observability and Tracingarize.com
- We need 3rd party Training-Run Assessments — LessWronglesswrong.com
- PostTrainBenchposttrainbench.com
- Concrete Evaluations to Investigate the OpenAI Model That Hacked Hugging Face — LessWronglesswrong.com
- confessions_paper.pdfcdn.openai.com
- The bitter lesson of LLM evalsparsed.com
- Training a Misaligned Reward Seekeralignment.anthropic.com