✳flâneur — a map of the web's best reading
On-Policy Distillation - Thinking Machines Lab
thinkingmachines.ai · 6,007 words · saved by 26 readers
On-policy, dense supervision is a useful tool for distillation
LLMs are capable of expert performance in focused domains, a result of several capabilities stacked together: perception of input, knowledge retrieval, plan selection, and reliable execution. This requires a stack of training approaches, which we can divide into three broad stages: Pre-training teaches general capacities such as language use, broad reasoning, and world knowledge. Mid-training imparts domain knowledge, such as code, medical databases, or internal company documents. Post-training elicits targeted behavior, such as instruction following, reasoning through math problems, or chat.
Explore this link on the map →saved by
- Elizabeth Qiu
- Kaylee George
- Rajan Agarwal
- Ratan Kaliani
- Emma Guo
- Asher P
- Uzay Girit
- Thu Than
- Aaron Pham
- Katherine Driscoll
- Dhruv Gautam
- Sudarsh K
related reading
- [2601.18734] Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Modelsarxiv.org
- SFT, RL, and On-Policy Distillation Through a Distributional Lens | whnrehiew.github.io
- [2604.00626] A Survey of On-Policy Distillation for Large Language Modelsarxiv.org
- Pedagogical RL: Teaching Models to Teach Themselves from Privileged Information - Noah Ziemsnoahziems.com
- [2605.10889] Unmasking On-Policy Distillation: Where It Helps, Where It Hurts, and Whyarxiv.org
- [2604.13010] Lightning OPD: Efficient Post-Training for Large Reasoning Models with Offline On-Policy Distillationarxiv.org
- Composer2.pdfcursor.com
- State of RL for reasoning LLMs | A. Weersaweers.de
- Self-Distillation Enables Continual Learningarxiv.org
- Bitter Lessons from Distillation Robustifies Unlearningbrucewlee.com
- will brown on X: "On SFT, RL, and on-policy distillation" / Xx.com
- [2305.18290] Direct Preference Optimization: Your Language Model is Secretly a Reward Modelarxiv.org