✳flâneur — a map of the web's best reading
Demystify Transformers: A Guide to Scaling Laws | by Yu-Cheng Tsai | Sage Ai | Medium
medium.com · 2,398 words · saved by 1 readers
Unpacking Transformer Technologies and Scaling Strategies
Demystify Transformers: A Guide to Scaling Laws Yu-Cheng Tsai 10 min read · Apr 30, 2024 -- 3 Listen Share Press enter or click to view image in full size This image was generated using DALL-E LLM Scaling Laws It’s no longer surprising that major cloud service providers and numerous companies are investing heavily in acquiring hundreds of thousands of GPU clusters and leveraging massive amounts of training data to develop Large Language Models (LLMs). Why are larger models considered better? When it comes to LLMs, the loss of next-token ( 1 token is about 0.75 of words) prediction is predictab
Explore this link on the map →saved by
related reading
- How To Scale Your Modeljax-ml.github.io
- Bits, FLOPS, and Watts: A Systems-Level Perspective of Scaling LLMs — Part 1 | by Asheesh Goja | Mediummedium.com
- New Scaling Laws for Large Language Models — LessWronglesswrong.com
- Scaling Laws, Carefully | Lil'Loglilianweng.github.io
- Transformer Math 101 | EleutherAI Blogblog.eleuther.ai
- Fermi estimate of future training runsdanieldewey.net
- LoRA Without Regret - Thinking Machines Labthinkingmachines.ai
- RL Scaling Laws for LLMs - by Cameron R. Wolfe, Ph.D.cameronrwolfe.substack.com
- The Scaling Hypothesis · Gwern.netgwern.net
- On neural scaling and the quanta hypothesisericjmichaud.com
- [2001.08361] Scaling Laws for Neural Language Modelsarxiv.org
- Reiner Pope – The math behind how LLMs are trained and serveddwarkesh.com