MiLo: Efficient Quantized MoE Inference with Mixture of Low-Rank Compensators
beichenhuang.github.io · 5,605 words · saved by 1 readers
N/A
M I L O : E FFICIENT Q UANTIZED M O E I NFERENCE WITH M IXTURE OF L OW-R ANK C OMPENSATORS Beichen Huang * 1 2 Yueming Yuan * 1 Zelei Shao * 1 Minjia Zhang 1 A BSTRACT A critical approach for efficiently deploying Mixture-of-Experts (MoE) models with massive parameters is quantization. However, state-of-the-art MoE models suffer from non-negligible accuracy loss with extreme quantization, such as under 4 bits. To address this, we introduce MiLo, a…
related reading
- Mixture of Experts Quantile Balancing: Validated at 32B-A5B (1e22 FLOPs) Scaleopenathena.ai
- LoRA Without Regret - Thinking Machines Labthinkingmachines.ai
- Mixture of Experts Explainedhuggingface.co
- A Guide to Quantization in LLMs | Symbl.aisymbl.ai
- Papers I’ve read this week, Mixture of Experts editionfinbarrtimbers.substack.com
- Quantization from the ground upngrok.com
- DeepSeek-V3: A Large-Scale MoE Pretraining Benchmark for MLPerf Training v6.0mlcommons.org
- The Big LLM Architecture Comparisonmagazine.sebastianraschka.com
- Better MoE model inference with warp decode · Cursorcursor.com
- Efficient LLM inferencefinbarrtimbers.substack.com
- Very Simple MoE Intro1a3orn.com
- Composer2.pdfcursor.com