flâneur — a map of the web's best reading

New Scaling Laws for Large Language Models - LessWrong

lesswrong.com · 4,016 words · saved by 1 readers

On March 29th, DeepMind published a paper, "Training Compute-Optimal Large Language Models", that shows that essentially everyone -- OpenAI, DeepMind, Microsoft, etc. -- has been training large langu…

x New Scaling Laws for Large Language Models — LessWrong GPT Language Models (LLMs) Machine Learning (ML) AI Frontpage 246 New Scaling Laws for Large Language Models by 1a3orn 1st Apr 2022 AI Alignment Forum 6 min read 22 246 Ω 71 On March 29th, DeepMind published a paper, "Training Compute-Optimal Large Language Models" , that shows that essentially everyone -- OpenAI, DeepMind, Microsoft, etc. -- has been training large language models with a deeply suboptimal use of compute. Following the new scaling laws that they propose for the optimal use of compute, DeepMind trains a new, 70-billion pa

Explore this link on the map →

related reading