flâneur

TPU Inference Externalization Full Steam Ahead - InferenceX

newsletter.semianalysis.com · 7,585 words · saved by 1 readers

InferenceX, Up to 50% Better Performance per Dollar, Rapid Externalization of TPU stack, Growing Customer Base, Ironwood, TPUv8i, Reducing CUDA Moat

For more than a decade, the industry has watched Google build an empire on its own silicon. Search, Ads, YouTube, and every generation of Gemini run on TPUs. Few accelerators have attracted as much architectural scrutiny or as much debate about what their performance and economics would look like outside the company that designed them. Anthropic being the biggest user of TPUs, surpassing Deepmind’s own use by 2029. Google’s internal success was never the question. The question was how much of that advantage the rest of the industry could actually get. Could you take an open-weight model,…

saved by

related reading