Bits, FLOPS, and Watts: A Systems-Level Perspective of Scaling LLMs — Part 1 | by Asheesh Goja | Medium
Scaling Large Language Models (LLMs) goes beyond simply increasing parameter counts. It emerges from a complex interplay of hardware, software, and algorithmic optimizations. These deep neural networks operate in high-dimensional vector spaces with complex loss landscapes and scaling power laws, where learning unfolds along intricate manifolds that often defy human intuition. Yet their existence is tethered to silicon and electricity — bound by thermodynamics, semiconductor physics, and the finite resources of our four-dimensional world. Ultimately, the mathematics of LLMs is bound by the very physics of silicon. (PDF version the paper) To understand their evolution, we must trace the journey of neural architectures from simple single-layer perceptrons to today’s trillion-parameter Transformers. At every stage, algorithmic breakthroughs and hardware advancements have propelled each other forward. As transistor scaling, memory bandwidth, and parallelism improved, each wave of innovation
Bits, FLOPS, and Watts: A Systems-Level Perspective of Scaling LLMs — Part 1 The Laws of Scaling Asheesh Goja 39 min read · Apr 5, 2025 -- Listen Share Introduction Scaling Large Language Models (LLMs) goes beyond simply increasing parameter counts. It emerges from a complex interplay of hardware, software, and algorithmic optimizations. These deep neural networks operate in high-dimensional vector spaces with complex loss landscapes and scaling power laws, where learning unfolds along intricate manifolds that often defy human intuition. Yet their existence is tethered to silicon and electrici
saved by
related reading
- How To Scale Your Modeljax-ml.github.io
- Demystify Transformers: A Guide to Scaling Laws | by Yu-Cheng Tsai | Sage Ai | Mediummedium.com
- Getting Caught Up to Modern LLM Research | Samarth Goeldev.samarthgoel.com
- Scaling Laws, Carefully | Lil'Loglilianweng.github.io
- Transformers Inference Optimization Toolset | AstraBlogastralord.github.io
- On neural scaling and the quanta hypothesisericjmichaud.com
- Reiner Pope – The math behind how LLMs are trained and serveddwarkesh.com
- LLM Resourcesforrestbicker.com
- How LLMs Actually Work | 0xkato0xkato.xyz
- The Big LLM Architecture Comparisonmagazine.sebastianraschka.com
- Scalable MatMul-free Language Modelingarxiv.org
- [2602.05970] Inverse Depth Scaling From Most Layers Being Similararxiv.org