Elicitation, the simplest way to understand post-training
interconnects.ai · 1,499 words · saved by 1 readers
An F1 analogy to help understand fast improvements in post-training on top of slow improvements in scaling.
Elicitation, the simplest way to understand post-training An F1 analogy to help understand fast improvements in post-training on top of slow improvements in scaling. Nathan Lambert Mar 10, 2025 48 9 Share Article voiceover 0:00 -8:25 Audio playback is not supported on your browser. Please upgrade. If you look at most of the models we've received from OpenAI, Anthropic, and Google in the last 18 months, you'll hear a lot of "Most of the improvements were in the post-training phase." The most recent one was Anthropic’s CEO Dario Amodei explaining Claude 3.7 on the Hard Fork Podcast : We are not
saved by
related reading
- The Extreme Inefficiency of RL for Frontier Models - Toby Ordtobyord.com
- [2606.07527] Post-training is (Massive) Supervised Learningarxiv.org
- aiesi_post-training_public.pdfkawine.github.io
- The Scaling Hypothesis · Gwern.netgwern.net
- Ilya Sutskever — We're moving from the age of scaling to the age of researchdwarkesh.com
- PostTrainBenchposttrainbench.com
- Why Does Reinforcement Learning Generalize? A Feature-Level Mechanistic Study of Post-Training in Large Language Modelsarxiv.org
- The upcoming GPT-3 moment for RL | Mechanize, Inc.mechanize.work
- Andrej Karpathy — AGI is still a decade awaydwarkesh.com
- Scaling is subtler than it seemsberen.io
- MAI-Thinking-1: Building a Hill-Climbing Machinemicrosoft.ai
- Early Data Exposure Improves Robustness to Subsequent Fine-Tuningarxiv.org