Demystify Transformers: A Guide to Scaling Laws | by Yu-Cheng Tsai | Sage Ai | Medium
medium.com · 2,398 words · saved by 1 readers
Unpacking Transformer Technologies and Scaling Strategies
Demystify Transformers: A Guide to Scaling Laws Yu-Cheng Tsai 10 min read · Apr 30, 2024 -- 3 Listen Share Press enter or click to view image in full size This image was generated using DALL-E LLM Scaling Laws It’s no longer surprising that major cloud service providers and numerous companies are investing heavily in acquiring hundreds of thousands of GPU clusters and leveraging massive amounts of training data to develop Large Language Models (LLMs). Why are larger models considered better? When it comes to LLMs, the loss of next-token ( 1 token is about 0.75 of words) prediction is predictab
saved by
related reading
- Scaling Laws, Carefully | Lil'Loglilianweng.github.io
- How To Scale Your Modeljax-ml.github.io
- Bits, FLOPS, and Watts: A Systems-Level Perspective of Scaling LLMs — Part 1 | by Asheesh Goja | Mediummedium.com
- Chinchillaarxiv.org
- New Scaling Laws for Large Language Models — LessWronglesswrong.com
- 2404.10102v1arxiv.org
- Fermi estimate of future training runsdanieldewey.net
- Transformer Math 101 | EleutherAI Blogblog.eleuther.ai
- IsoFLOP curves of large language models are extremely flatseverelytheoretical.wordpress.com
- [Jan 7 2026] nanochat miniseries v1 · karpathy nanochat · Discussion #420github.com
- The Scaling Hypothesis · Gwern.netgwern.net
- RL Scaling Laws for LLMs - by Cameron R. Wolfe, Ph.D.cameronrwolfe.substack.com