Scaling Laws, Carefully | Lil'Log
Scaling laws are one of the most critical empirical findings in deep learning. The observation is simple in form: the training loss $L$ decreases predictably as we scale up model size $N$, dataset size $D$, and compute $C$, following a power-law curve, which appears as a straight line on a log-log plot. We can view scaling laws as a framework for describing the relationship between compute, loss, model size and data; at its core, it is about how to allocate precious compute optimally between $N$ and $D$.
Table of Contents Early days: ML loss predictability Scaling Laws in Data-Infinite Region Kaplan et al.'s Scaling Laws Chinchilla Scaling Laws Method 1: Fix model sizes, vary the token budget Method 2: IsoFLOP profiles Method 3: Parametric fit Reconciling Kaplan and Chinchilla Why power law? Scaling Laws in Data-Limited Region Trickiness of Fitting Scaling Laws in Reality Toy simulation Citation References Scaling laws are one of the most critical empirical findings in deep learning. The observation is simple in form: the training loss $L$ decreases predictably as we scale up model size $N$, d
Explore this link on the map →related reading
- New Scaling Laws for Large Language Models — LessWronglesswrong.com
- Demystify Transformers: A Guide to Scaling Laws | by Yu-Cheng Tsai | Sage Ai | Mediummedium.com
- On neural scaling and the quanta hypothesisericjmichaud.com
- The Scaling Hypothesis · Gwern.netgwern.net
- [2001.08361] Scaling Laws for Neural Language Modelsarxiv.org
- Scaling laws literature review | Epoch AIepochai.org
- The Little Book of Deep Learningfleuret.org
- RL Scaling Laws for LLMs - by Cameron R. Wolfe, Ph.D.cameronrwolfe.substack.com
- Fermi estimate of future training runsdanieldewey.net
- chinchilla's wild implications — AI Alignment Forumalignmentforum.org
- [2602.05970] Inverse Depth Scaling From Most Layers Being Similararxiv.org
- [2510.03280] Training Optimal Large Diffusion Language Modelsar5iv.labs.arxiv.org