Distillation Walkthrough
vladfeinberg.com · 1,446 words · saved by 2 readers
Vlad's Blog
Distillation Walkthrough Distillation is a critical technique towards improving a network’s quality while keeping its serving latency constant. This is becoming crucial as people focus on serving larger and larger models. Image found on LinkedIn . Distillation is a powerful technique, but a couple things about it are quite mystical. The purpose of this post is to: Provide a very high level explainer (but mostly refer to source papers) of distillation. Show that you can create a simple, linear example where distillation works. How is distillation implemented? At an algorithmic level, distillati
saved by
related reading
- On-Policy Distillation - Thinking Machines Labthinkingmachines.ai
- [1511.03643] Unifying distillation and privileged informationarxiv.org
- SFT, RL, and On-Policy Distillation Through a Distributional Lens | whnrehiew.github.io
- What is Model Distillation?labelbox.com
- Nitrobrew: Fast, Lossless Distillation for Free | Tildeblog.tilderesearch.com
- Everything You Need to Know about Knowledge Distillationhuggingface.co
- [2605.23857] Strong Teacher Not Needed? On Distillation in LLM Pretrainingarxiv.org
- Zero-Shot Knowledge Distillation in Deep Networksarxiv.org
- [2605.10889] Unmasking On-Policy Distillation: Where It Helps, Where It Hurts, and Whyarxiv.org
- The Little Book of Deep Learningfleuret.org
- [2604.00626] A Survey of On-Policy Distillation for Large Language Modelsarxiv.org
- [1503.02531] Distilling the Knowledge in a Neural Networkarxiv.org