✳flâneur — a map of the web's best reading
What is Model Distillation?
labelbox.com · 1,705 words · saved by 1 readers
Model distillation compresses large models for deployment, transferring knowledge from a teacher to a smaller student model efficiently.
Explore this link on the map →saved by
related reading
- On-Policy Distillation - Thinking Machines Labthinkingmachines.ai
- [2604.00626] A Survey of On-Policy Distillation for Large Language Modelsarxiv.org
- Distillation Walkthroughvladfeinberg.com
- Nitrobrew: Fast, Lossless Distillation for Free | Tildeblog.tilderesearch.com
- [2601.18734] Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Modelsarxiv.org
- [2605.23857] Strong Teacher Not Needed? On Distillation in LLM Pretrainingarxiv.org
- [1503.02531] Distilling the Knowledge in a Neural Networkarxiv.org
- Bitter Lessons from Distillation Robustifies Unlearningbrucewlee.com
- [2604.13010] Lightning OPD: Efficient Post-Training for Large Reasoning Models with Offline On-Policy Distillationarxiv.org
- GitHub - jacobhilton/deep_learning_curriculum: Language model alignment-focused deep learning curriculum · GitHubgithub.com
- [2605.10889] Unmasking On-Policy Distillation: Where It Helps, Where It Hurts, and Whyarxiv.org
- AI Model & API Providers Analysis | Artificial Analysisartificialanalysis.ai