Zero-Shot Knowledge Distillation in Deep Networks
arxiv.org · 6,659 words · saved by 1 readers
N/A
Zero-Shot Knowledge Distillation in Deep Networks Gaurav Kumar Nayak * 1 Konda Reddy Mopuri * 2 Vaisakh Shaj * 3 R. Venkatesh Babu 1 Anirban Chakraborty 1 Abstract limited resource environments or when real-time inference Knowledge distillation deals with the problem of is expected. On the…
saved by
related reading
- Everything You Need to Know about Knowledge Distillationhuggingface.co
- On-Policy Distillation - Thinking Machines Labthinkingmachines.ai
- What is Model Distillation?labelbox.com
- Distillation Walkthroughvladfeinberg.com
- [2605.23857] Strong Teacher Not Needed? On Distillation in LLM Pretrainingarxiv.org
- [1511.03643] Unifying distillation and privileged informationarxiv.org
- [2604.00626] A Survey of On-Policy Distillation for Large Language Modelsarxiv.org
- LLM Optimization via Synthetic Distillationanarchyai.substack.com
- Self-Distillation Enables Continual Learningarxiv.org
- [2601.18734] Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Modelsarxiv.org
- [1503.02531] Distilling the Knowledge in a Neural Networkarxiv.org
- Self-Distillation Enables Continual Learningarxiv.org