Redesigning the Inference Chip: From Nvidia GPU's Flaws to OpenAI Jalapeño
zartbot.github.io · 9,714 words · saved by 1 readers
First principles · Core Slice architecture · Memory subsystem · On-chip network · Software stack & design loop
TL;DR If I were reborn as the Infra lead of some foundation-model team, I would certainly end up with all kinds of complaints about Nvidia's GPUs and the system as a whole, and then go build my own chip... For Nvidia's networking part there is nothing to discuss: RoCE has all sorts of defects, and the DPU also comes with a pile of performance and security problems. No wonder Nvidia has started hyping Scale-in again this year, but for a team that has never run a cloud, it will take at least 5 more years to get mature in this area. As an example, we recently did something in our MaaS online…
saved by
related reading
- OpenAI Jalapeño: Better Than Nvidia Blackwellnewsletter.semianalysis.com
- Meta harness makes 10 times better Kimi K3 chipluoluo.ai
- The Short Case for Nvidia Stock | YouTube Transcript Optimizeryoutubetranscriptoptimizer.com
- AI Chip Architecturesjacobpeake.com
- How to Think About GPUs | How To Scale Your Modeljax-ml.github.io
- AI Chip Architecturesjepeake.com
- Reiner Pope – The math behind how LLMs are trained and serveddwarkesh.com
- The Best GPUs for Deep Learning in 2023 — An In-depth Analysistimdettmers.com
- An Interview with MatX CEO Reiner Pope About LLM Chipschipstrat.com
- Inside NVIDIA GPUs: Anatomy of high performance matmul kernels - Aleksa Gordićaleksagordic.com
- Dissecting the NVIDIA Volta GPU Architecture via Microbenchmarkingarxiv.org
- The Architecture of Dominance: NVIDIA’s Rubin CPX and the $254 Billion Inference Warsshanakaanslemperera.substack.com