Mixture of Experts Quantile Balancing: Validated at 32B-A5B (1e22 FLOPs) Scale | Open Athena
Quantile Balancing (QB) is a hyperparameter-free load balancer for Mixture of Experts models, introduced by Jianlin Su. We validated it on a 32B-A5B (1e22 FLOPs) Marin run over 326B tokens: zero...
This post highlights Quantile Balancing (QB), a load balancing method for Mixture of Experts (MoE) models developed by Jianlin Su. The Marin team validated QB on a 32B total 5B active-parameter (32B-A5B) model trained on Marin infrastructure over 326B tokens (1e22 FLOP compute budget). In this post, I will get into our motivation for using the MoE model family, explain why load balancing is necessary, survey prior approaches to balancing, and share performance results for QB. Summary QB is a hyperparameter-free MoE load balancer from Jianlin Su. The Marin team validated it at 32B-A5B /…
saved by
related reading
- DeepSeek-V3: A Large-Scale MoE Pretraining Benchmark for MLPerf Training v6.0mlcommons.org
- [2607.16051] Loop the Loopies!arxiv.org
- Mixture-of-Kittens: our open-source MoE megakernel for NVL72scursor.com
- Papers I’ve read this week, Mixture of Experts editionfinbarrtimbers.substack.com
- Better MoE model inference with warp decode · Cursorcursor.com
- Mixture of experts - Wikipediaen.wikipedia.org
- Mixture of Experts Explainedhuggingface.co
- MiLo: Efficient Quantized MoE Inference with Mixture of Low-Rank Compensatorsbeichenhuang.github.io
- Very Simple MoE Intro1a3orn.com
- [2202.08906] ST-MoE: Designing Stable and Transferable Sparse Expert Modelsar5iv.labs.arxiv.org
- frontier model training methodologies | Alex Wa's Blogdjdumpling.github.io
- Mixtral of Expertsarxiv.org