flâneur — a map of the web's best reading

More Efficient In-Context Learning with GLaM – Google Research Blog

blog.research.google · 1,762 words · saved by 1 readers

Large language models (e.g., GPT-3) have many significant capabilities, such as performing few-shot learning across a wide array of tasks, including reading comprehension and question answering with very few or no training examples. While these models can perform better by simply using more parameters, training and serving these large models can be very computationally intensive. Is it possible to train and use these models more efficiently? In “GLaM: Efficient Scaling of Language Models with Mixture-of-Experts”, we introduce the Generalist Language Model (GLaM), a trillion weight model that can be trained and served efficiently (in terms of computation and energy use) thanks to sparsity, and achieves competitive performance on multiple few-shot learning tasks. GLaM’s performance compares favorably to a dense language model, GPT-3 (175B) with significantly improved learning efficiency across 29 public NLP benchmarks in seven categories, spanning language completion, open-domain questio

More Efficient In-Context Learning with GLaM Skip to main content More Efficient In-Context Learning with GLaM December 9, 2021 Posted by Andrew M Dai and Nan Du, Research Scientists, Google Research, Brain Team Quick links Share Copy link × Large language models (e.g., GPT-3 ) have many significant capabilities, such as performing few-shot learning across a wide array of tasks, including reading comprehension and question answering with very few or no training examples. While these models can perform better by simply using more parameters, training and serving these large models can be very c

Explore this link on the map →

saved by

related reading