flâneur — a map of the web's best reading

Go smol or go home | Harm de Vries

harmdevries.com · 9,051 words · saved by 3 readers

The Chinchilla scaling laws suggest we haven’t reached the limit of training smaller models for longer.

Go smol or go home | Harm de Vries Search Go smol or go home Why we should train smaller LLMs on more tokens Harm de Vries Last updated on Jul 3, 2023 110 min read If you have access to a big compute cluster and are planning to train a Large Language Model (LLM), you will need to make a decision on how to allocate your compute budget. This involves selecting the number of model parameters $N$ and the number of training tokens $D$. By applying the scaling laws , you can get guidance on how to reach the best model performance for your given compute budget, and find the optimal distribution of co

Explore this link on the map →

saved by

related reading