✳flâneur — a map of the web's best reading
Elicitation, the simplest way to understand post-training
interconnects.ai · 1,499 words · saved by 1 readers
An F1 analogy to help understand fast improvements in post-training on top of slow improvements in scaling.
Elicitation, the simplest way to understand post-training An F1 analogy to help understand fast improvements in post-training on top of slow improvements in scaling. Nathan Lambert Mar 10, 2025 48 9 Share Article voiceover 0:00 -8:25 Audio playback is not supported on your browser. Please upgrade. If you look at most of the models we've received from OpenAI, Anthropic, and Google in the last 18 months, you'll hear a lot of "Most of the improvements were in the post-training phase." The most recent one was Anthropic’s CEO Dario Amodei explaining Claude 3.7 on the Hard Fork Podcast : We are not
Explore this link on the map →saved by
related reading
- The Scaling Hypothesis · Gwern.netgwern.net
- RLHF & Post-Training Course by Nathan Lambertrlhfbook.com
- Andrej Karpathy — AGI is still a decade awaydwarkesh.com
- RLHF: Reinforcement Learning from Human Feedbackhuyenchip.com
- PostTrainBenchposttrainbench.com
- [2606.12360] Anatomy of Post-Training: Using Interpretability to Characterize Data and Shape the Learning Signalarxiv.org
- rlhfbook.com/book.pdfrlhfbook.com
- Pre, Mid, Post-Training Way of Life - by Tina Hefakepixels.substack.com
- MAI-Thinking-1: Building a Hill-Climbing Machinemicrosoft.ai
- RL Scaling Laws for LLMs - by Cameron R. Wolfe, Ph.D.cameronrwolfe.substack.com
- DeepSeek R1's recipe to replicate o1 and the future of reasoning LMsinterconnects.ai
- AI capabilities can be significantly improved without expensive retraining | Epoch AIepoch.ai