✳flâneur — a map of the web's best reading
The Architecture of Dominance: NVIDIA’s Rubin CPX and the $254 Billion Inference Wars
shanakaanslemperera.substack.com · 2,783 words · saved by 1 readers
How a chip designed to process context—not generate it—may determine who controls the economics of artificial intelligence through 2030
The Architecture of Dominance: NVIDIA’s Rubin CPX and the $254 Billion Inference Wars How a chip designed to process context—not generate it—may determine who controls the economics of artificial intelligence through 2030 Shanaka Anslem Perera Dec 30, 2025 ∙ Paid 9 2 Share Shanaka Anslem Perera December 30, 2025 I. The Confession Hidden in Plain Sight On September 9, 2025, NVIDIA unveiled a chip that shouldn’t exist. The Rubin CPX contains no High-Bandwidth Memory. It uses GDDR7—the same commodity memory found in gaming cards. Its die is monolithic, avoiding the exotic CoWoS packaging that def
Explore this link on the map →related reading
- The Short Case for Nvidia Stock | YouTube Transcript Optimizeryoutubetranscriptoptimizer.com
- Another Giant Leap: The Rubin CPX Specialized Accelerator & Rack – SemiAnalysissemianalysis.com
- Transformer Inference Arithmetic | kipply's blogkipp.ly
- The Inference Shift – Stratechery by Ben Thompsonstratechery.com
- Reiner Pope – The math behind how LLMs are trained and serveddwarkesh.com
- An Interview with MatX CEO Reiner Pope About LLM Chipschipstrat.com
- Real-time LLM Inference on Standard Datacenter GPUs (3,000 tokens/s per request)blog.kog.ai
- LLM Inference Performance Engineering: Best Practices | Databricks Blogdatabricks.com
- How to Think About GPUs | How To Scale Your Modeljax-ml.github.io
- Best practices to accelerate inference for large-scale production workloadstogether.ai
- Groq Inference Tokenomics: Speed, But At What Cost?semianalysis.com
- My picture of the present in AI — LessWronglesswrong.com