ST-MoE: Designing Stable and Transferable Sparse Expert Models | HTML5
Scale has opened new frontiers in natural language processing – but at a high cost. In response, Mixture-of-Experts (MoE) and Switch Transformers have been proposed as an energy efficient path to even larger and more capable language models. But advancing the state-of-the-art across a broad set of natural language tasks has been hindered by training instabilities and uncertain quality during fine-tuning. Our work focuses on these issues and acts as a design guide. We conclude by scaling a sparse model to 269B parameters, with a computational cost comparable to a 32B dense encoder-decoder Transformer (Stable and Transferable Mixture-of-Experts or ST-MoE-32B). For the first time, a sparse model achieves state-of-the-art performance in transfer learning, across a diverse set of tasks including reasoning (SuperGLUE, ARC Easy, ARC Challenge), summarization (XSum, CNN-DM), closed book question answering (WebQA, Natural Questions), and adversarially constructed tasks (Winogrande, ANLI R3). 1
\UseRawInputEncoding ST-MoE: Designing Stable and Transferable Sparse Expert Models Barret Zoph Google Brain &Irwan Bello 1 1 footnotemark: 1 Google Brain \AND Sameer Kumar Google &Nan Du Google Brain &Yanping Huang Google Brain \AND Jeff Dean Google Research &Noam Shazeer 2 2 footnotemark: 2 Google Brain &William Fedus 1 1 footnotemark: 1 Google Brain Equal contribution. Correspondence to { barretzoph,liamfedus}@google.com .Work was done while at Google. Abstract Scale has opened new frontiers in natural language processing – but at a high cost. In response, Mixture-of-Experts (MoE) and Switc
Explore this link on the map →related reading
- Papers I’ve read this week, Mixture of Experts editionfinbarrtimbers.substack.com
- Very Simple MoE Intro1a3orn.com
- frontier model training methodologies | Alex Wa's Blogdjdumpling.github.io
- LoRA Without Regret - Thinking Machines Labthinkingmachines.ai
- More Efficient In-Context Learning with GLaMblog.research.google
- [2209.01667] A Review of Sparse Expert Models in Deep Learningarxiv.org
- Composer2.pdfcursor.com
- Mixture of experts - Wikipediaen.wikipedia.org
- Monet: Mixture of Monosemantic Experts for Transformers Explained — LessWronglesswrong.com
- [1701.06538] Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layerarxiv.org
- On the Tradeoffs of SSMs and Transformers | Goomba Labgoombalab.github.io
- An Alternative to Test-Time Scalingrentry.org