✳flâneur — a map of the web's best reading
Go smol or go home | Harm de Vries
harmdevries.com · 9,051 words · saved by 3 readers
The Chinchilla scaling laws suggest we haven’t reached the limit of training smaller models for longer.
Go smol or go home | Harm de Vries Search Go smol or go home Why we should train smaller LLMs on more tokens Harm de Vries Last updated on Jul 3, 2023 110 min read If you have access to a big compute cluster and are planning to train a Large Language Model (LLM), you will need to make a decision on how to allocate your compute budget. This involves selecting the number of model parameters $N$ and the number of training tokens $D$. By applying the scaling laws , you can get guidance on how to reach the best model performance for your given compute budget, and find the optimal distribution of co
Explore this link on the map →saved by
related reading
- GitHub - inverse-scaling/prize: A prize for finding tasks that cause large language models to show inverse scaling · GitHubgithub.com
- On neural scaling and the quanta hypothesisericjmichaud.com
- GitHub - google-research/tuning_playbook: A playbook for systematically maximizing the performance of deep learning models. · GitHubgithub.com
- GitHub - jacobhilton/deep_learning_curriculum: Language model alignment-focused deep learning curriculum · GitHubgithub.com
- Scaling Laws, Carefully | Lil'Loglilianweng.github.io
- GLaM: Efficient Scaling of Language Models with Mixture-of-Expertsarxiv.org
- chinchilla's wild implications — AI Alignment Forumalignmentforum.org
- GitHub - karpathy/nanochat: The best ChatGPT that $100 can buy. · GitHubgithub.com
- GitHub - openai/parameter-golf: Train the smallest LM you can that fits in 16MB. Best model wins! · GitHubgithub.com
- MatX: High-throughput chips for LLMsmatx.com
- Neuronpedianeuronpedia.org
- Training a compute-optimal gpt2-small – Tomek Korbak — personal homepagetomekkorbak.com