ST-MoE: Designing Stable and Transferable Sparse Expert Models | HTML5
Scale has opened new frontiers in natural language processing – but at a high cost. In response, Mixture-of-Experts (MoE) and Switch Transformers have been proposed as an energy efficient path to even larger and more capable language models. But advancing the state-of-the-art across a broad set of natural language tasks has been hindered by training instabilities and uncertain quality during fine-tuning. Our work focuses on these issues and acts as a design guide. We conclude by scaling a sparse model to 269B parameters, with a computational cost comparable to a 32B dense encoder-decoder Transformer (Stable and Transferable Mixture-of-Experts or ST-MoE-32B). For the first time, a sparse model achieves state-of-the-art performance in transfer learning, across a diverse set of tasks including reasoning (SuperGLUE, ARC Easy, ARC Challenge), summarization (XSum, CNN-DM), closed book question answering (WebQA, Natural Questions), and adversarially constructed tasks (Winogrande, ANLI R3). 1
\UseRawInputEncoding ST-MoE: Designing Stable and Transferable Sparse Expert Models Barret Zoph Google Brain &Irwan Bello 1 1 footnotemark: 1 Google Brain \AND Sameer Kumar Google &Nan Du Google Brain &Yanping Huang Google Brain \AND Jeff Dean Google Research &Noam Shazeer 2 2 footnotemark: 2 Google Brain &William Fedus 1 1 footnotemark: 1 Google Brain Equal contribution. Correspondence to { barretzoph,liamfedus}@google.com .Work was done while at Google. Abstract Scale has opened new frontiers in natural language processing – but at a high cost. In response, Mixture-of-Experts (MoE) and Switc
related reading
- [2101.03961] Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsityarxiv.org
- Papers I’ve read this week, Mixture of Experts editionfinbarrtimbers.substack.com
- Mixture of Experts Explainedhuggingface.co
- [2209.01667] A Review of Sparse Expert Models in Deep Learningarxiv.org
- Very Simple MoE Intro1a3orn.com
- frontier model training methodologies | Alex Wa's Blogdjdumpling.github.io
- [1701.06538] Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layerarxiv.org
- Mixtral of Expertsarxiv.org
- LoRA Without Regret - Thinking Machines Labthinkingmachines.ai
- More Efficient In-Context Learning with GLaMblog.research.google
- Composer2.pdfcursor.com
- DeepSeek-V3: A Large-Scale MoE Pretraining Benchmark for MLPerf Training v6.0mlcommons.org