✳flâneur — a map of the web's best reading
Distillation Walkthrough
vladfeinberg.com · 1,446 words · saved by 1 readers
Vlad's Blog
Distillation Walkthrough Distillation is a critical technique towards improving a network’s quality while keeping its serving latency constant. This is becoming crucial as people focus on serving larger and larger models. Image found on LinkedIn . Distillation is a powerful technique, but a couple things about it are quite mystical. The purpose of this post is to: Provide a very high level explainer (but mostly refer to source papers) of distillation. Show that you can create a simple, linear example where distillation works. How is distillation implemented? At an algorithmic level, distillati
Explore this link on the map →saved by
related reading
- On-Policy Distillation - Thinking Machines Labthinkingmachines.ai
- [1511.03643] Unifying distillation and privileged informationarxiv.org
- Nitrobrew: Fast, Lossless Distillation for Free | Tildeblog.tilderesearch.com
- [2605.23857] Strong Teacher Not Needed? On Distillation in LLM Pretrainingarxiv.org
- [2605.10889] Unmasking On-Policy Distillation: Where It Helps, Where It Hurts, and Whyarxiv.org
- [2604.00626] A Survey of On-Policy Distillation for Large Language Modelsarxiv.org
- What is Model Distillation?labelbox.com
- SFT, RL, and On-Policy Distillation Through a Distributional Lens | whnrehiew.github.io
- [2601.18734] Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Modelsarxiv.org
- AI safety techniques leveraging distillation — LessWronglesswrong.com
- will brown on X: "On SFT, RL, and on-policy distillation" / Xx.com
- Bitter Lessons from Distillation Robustifies Unlearningbrucewlee.com