[2606.07527] Post-training is (Massive) Supervised Learning
Abstract:The prevailing paradigm for training LLMs has evolved to rely on a massive post-training phase consisting of SFT and RL. In this position paper, we argue that this methodology effectively marks a reversion to the ``pre-train then fine-tune'' approach of the BERT era, explicitly tailoring models to the desired behaviors and specific benchmarks on which they are evaluated. We begin with a historical overview of LLMs, describing the different phases of the LLM evolution. We argue that the current landscape is remarkably similar to the early days of LLMs, where task performance heavily relied on fitting the models to in-distribution datasets. To empirically demonstrate this, we compare pre-trained models to randomly initialized ones, by fine-tuning both variants on modern reasoning datasets and evaluating them on competitive math and code benchmarks. We show that models post-trained from scratch yield highly non-trivial performance. Our findings suggest that current post-training methodologies function primarily as a distribution-fitting mechanism. We finish by positing that developing generally capable models and systems requires moving beyond extensive post-training for predefined behaviors, shifting instead toward training procedures where models ``learn how to learn''.
Post-training is (Massive) Supervised Learning Michael Hassid1,2 , Yossi Adi1,2 , Roy Schwartz2 1 FAIR, Meta AI 2 The Hebrew University of Jerusalem michael.hassid@mail.huji.ac.il Abstract arXiv:2606.07527v1 [cs.CL] 20 Apr 2026 The prevailing paradigm for training LLMs has evolved to rely on a massive…
saved by
related reading
- LLM Post-Training: A Deep Dive into Reasoning Large Language Modelsarxiv.org
- Why Does Reinforcement Learning Generalize? A Feature-Level Mechanistic Study of Post-Training in Large Language Modelsarxiv.org
- GenAI Handbookgenai-handbook.github.io
- PostTrainBenchposttrainbench.com
- [2602.05910] Chunky Post-Training: Data Driven Failures of Generalizationarxiv.org
- The Extreme Inefficiency of RL for Frontier Models - Toby Ordtobyord.com
- [2005.14165] Language Models are Few-Shot Learnersarxiv.org
- Getting Caught Up to Modern LLM Research | Samarth Goeldev.samarthgoel.com
- aiesi_post-training_public.pdfkawine.github.io
- Elicitation, the simplest way to understand post-traininginterconnects.ai
- MAI-Thinking-1: Building a Hill-Climbing Machinemicrosoft.ai
- Recent Advances in Language Model Fine-tuningruder.io