IsoFLOP curves of large language models are extremely flat – Severely Theoretical
An interesting detail in the recently released Llama-3 technical report has caught my eye (p. 8): This has caught my eye, since I had noted the same phenomenon in a previous post about the Chinchil…
An interesting detail in the recently released Llama-3 technical report has caught my eye (p. 8): This has caught my eye, since I had noted the same phenomenon in a previous post about the Chinchilla scaling laws (more than two years ago) to argue for training smaller models (point 4 in that post). I’m glad that this observation is finally being taken seriously, but I think the quotation above from the Llama-3 paper still underestimates the extent of this isoFLOP flatness issue. The performance of these models is not just robust to small variations in model size around the optimal, but it…
saved by
related reading
- Scaling Laws, Carefully | Lil'Loglilianweng.github.io
- Chinchillaarxiv.org
- Demystify Transformers: A Guide to Scaling Laws | by Yu-Cheng Tsai | Sage Ai | Mediummedium.com
- LoRA Without Regret - Thinking Machines Labthinkingmachines.ai
- How To Scale Your Modeljax-ml.github.io
- Fermi estimate of future training runsdanieldewey.net
- [Jan 7 2026] nanochat miniseries v1 · karpathy nanochat · Discussion #420github.com
- Scaling Laws That Extrapolate 300× Past the Fitopenathena.ai
- New Scaling Laws for Large Language Models — LessWronglesswrong.com
- 2404.10102v1arxiv.org
- frontier model training methodologies | Alex Wa's Blogdjdumpling.github.io
- Mosaic LLMs: GPT-3 quality formosaicml.com