Another Giant Leap: The Rubin CPX Specialized Accelerator & Rack – SemiAnalysis
Nvidia announced the Rubin CPX, a solution that is specifically designed to be optimized for the prefill phase, with the single-die Rubin CPX heavily emphasizing compute FLOPS over memory bandwidth. This is a game changer for inference, and its significance is surpassed only by the March 2024 announcement of the GB200 NVL72 Oberon rack-scale form factor. Only with hardware specialized to the very different phases of inference, prefill and decode, can disaggregated serving achieve its full potential. As a result, the rack system design gap between Nvidia and its competitors has become canyon-sized. AMD and custom silicon competitors may have made a small step forward in emulating Nvidia’s 72-GPU rack scale design, but Nvidia has just made another Giant Leap, again leaving competitors very distant objects in the rear-view mirror. AMD and ASIC providers have already been investing heavily to catch up in terms of their own rack-scale solutions. AMD in particular has been working tirelessly
Another Giant Leap: The Rubin CPX Specialized Accelerator & Rack – SemiAnalysis Skip to content September 10, 2025 Another Giant Leap: The Rubin CPX Specialized Accelerator & Rack New Prefill Specialized GPU, Rack Architecture, BOM, Disaggregated PD, Higher Perf per TCO, Lower TCO, GDDR7 & HBM Market Trends 23 minutes 2 comments on Another Giant Leap: The Rubin CPX Specialized Accelerator & Rack By Dylan Patel , Daniel Nishball , Kimbo Chen , Myron Xie , Wega Chu , Gerald Wong , Kang Wen Cheang and Ivan Chiam Share using Native tools Share Copied to clipboard Share on LinkedIn (Opens in
Explore this link on the map →saved by
related reading
- NVIDIA Tensor Core Evolution: From Volta To Blackwellnewsletter.semianalysis.com
- H100 vs GB200 NVL72 Training Benchmarks – Power, TCO, and Reliability Analysis, Software Improvement Over Time – SemiAnalysissemianalysis.com
- The Architecture of Dominance: NVIDIA’s Rubin CPX and the $254 Billion Inference Warsshanakaanslemperera.substack.com
- How to Think About TPUs | How To Scale Your Modeljax-ml.github.io
- Inside NVIDIA GPUs: Anatomy of high performance matmul kernels - Aleksa Gordićaleksagordic.com
- All About Rooflines | How To Scale Your Modeljax-ml.github.io
- The Short Case for Nvidia Stock | YouTube Transcript Optimizeryoutubetranscriptoptimizer.com
- Inside NVIDIA Rubin GPU Architecture: Powering the Era of Agentic AI | NVIDIA Technical Blogdeveloper.nvidia.com
- Transformer Inference Arithmetic | kipply's blogkipp.ly
- The Best GPUs for Deep Learning in 2023 — An In-depth Analysistimdettmers.com
- How to Think About GPUs | How To Scale Your Modeljax-ml.github.io
- The Inference Shift – Stratechery by Ben Thompsonstratechery.com