flâneur — a map of the web's best reading

Scaling Laws, Carefully | Lil'Log

lilianweng.github.io · 5,203 words · saved by 1 readers

Scaling laws are one of the most critical empirical findings in deep learning. The observation is simple in form: the training loss $L$ decreases predictably as we scale up model size $N$, dataset size $D$, and compute $C$, following a power-law curve, which appears as a straight line on a log-log plot. We can view scaling laws as a framework for describing the relationship between compute, loss, model size and data; at its core, it is about how to allocate precious compute optimally between $N$ and $D$.

Table of Contents Early days: ML loss predictability Scaling Laws in Data-Infinite Region Kaplan et al.'s Scaling Laws Chinchilla Scaling Laws Method 1: Fix model sizes, vary the token budget Method 2: IsoFLOP profiles Method 3: Parametric fit Reconciling Kaplan and Chinchilla Why power law? Scaling Laws in Data-Limited Region Trickiness of Fitting Scaling Laws in Reality Toy simulation Citation References Scaling laws are one of the most critical empirical findings in deep learning. The observation is simple in form: the training loss $L$ decreases predictably as we scale up model size $N$, d

Explore this link on the map →

related reading