An empirical analysis of compute-optimal large language model training
We investigate the optimal model and dataset size for training a transformer language model under a given compute budget. We find that current large language models are significantly undertrained, a consequence of the recent focus on scaling language models whilst keeping the amount of training data constant. By training \nummodels language models ranging from 70 million to 10 billion parameters on 5 to 500 billion tokens, we find that for compute-optimal training, the model size and the training dataset size should be scaled equally: for every doubling of model size the training dataset size should also be doubled. We test this hypothesis by training a more compute-optimal model, \Chinchilla, using the same compute budget as \gopher but with 70B parameters and 4$\times$ more data. \chinchilla uniformly and significantly outperforms \Gopher, GPT-3, Jurassic-1, and \mtnlg on a large range of downstream evaluation tasks. As a highlight, \chinchilla reaches an average accuracy of 67.5\% on the MMLU benchmark, over a 7\% improvement over \gopher.
Publications — Google DeepMind Skip to main content Publications Explore a selection of our recent research on some of the most complex and interesting challenges in AI. 259 publications 10 July 2026 Quantifying the Salience of Geo-Cultural Values for Pluralistic Safety Alignment 6 July 2026 The Case for Globally Beneficial Technology 2 July 2026 Towards Structural Understanding of LLM Overthinking 26 June 2026 Bridging the Scale Gap: Augmenting Human Red-Teaming to Uncover Latent Risks in T2I Models 26 June 2026 Real-Time Group Dynamics with LLM Facilitation: Evidence from a Charity Allocatio
Explore this link on the map →related reading
- [2203.15556] Training Compute-Optimal Large Language Modelsarxiv.org
- AI in 2025: gestalt — LessWronglesswrong.com
- Composer2.pdfcursor.com
- Large Language Models Reading List | Sebastian Raschka, PhDsebastianraschka.com
- Automated Alignment Researchers: Using large language models to scale scalable oversight \ Anthropicanthropic.com
- Demystify Transformers: A Guide to Scaling Laws | by Yu-Cheng Tsai | Sage Ai | Mediummedium.com
- IsoCompute Playbook: Optimally Scaling Sampling Compute for RL Training of LLMscompute-optimal-rl-llm-scaling.github.io
- LLM Resourcesforrestbicker.com
- New Scaling Laws for Large Language Models — LessWronglesswrong.com
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- Training a compute-optimal gpt2-small – Tomek Korbak — personal homepagetomekkorbak.com
- Thoughts on the Alignment Implications of Scaling Language Models | Leo Gaobmk.sh