flâneur — a map of the web's best reading

ST-MoE: Designing Stable and Transferable Sparse Expert Models | HTML5

ar5iv.labs.arxiv.org · 20,364 words · saved by 1 readers

Scale has opened new frontiers in natural language processing – but at a high cost. In response, Mixture-of-Experts (MoE) and Switch Transformers have been proposed as an energy efficient path to even larger and more capable language models. But advancing the state-of-the-art across a broad set of natural language tasks has been hindered by training instabilities and uncertain quality during fine-tuning. Our work focuses on these issues and acts as a design guide. We conclude by scaling a sparse model to 269B parameters, with a computational cost comparable to a 32B dense encoder-decoder Transformer (Stable and Transferable Mixture-of-Experts or ST-MoE-32B). For the first time, a sparse model achieves state-of-the-art performance in transfer learning, across a diverse set of tasks including reasoning (SuperGLUE, ARC Easy, ARC Challenge), summarization (XSum, CNN-DM), closed book question answering (WebQA, Natural Questions), and adversarially constructed tasks (Winogrande, ANLI R3). 1

\UseRawInputEncoding ST-MoE: Designing Stable and Transferable Sparse Expert Models Barret Zoph Google Brain &Irwan Bello 1 1 footnotemark: 1 Google Brain \AND Sameer Kumar Google &Nan Du Google Brain &Yanping Huang Google Brain \AND Jeff Dean Google Research &Noam Shazeer 2 2 footnotemark: 2 Google Brain &William Fedus 1 1 footnotemark: 1 Google Brain Equal contribution. Correspondence to { barretzoph,liamfedus}@google.com .Work was done while at Google. Abstract Scale has opened new frontiers in natural language processing – but at a high cost. In response, Mixture-of-Experts (MoE) and Switc

Explore this link on the map →

related reading