flâneur — a map of the web's best reading

Multi-Datacenter Training: OpenAI's Ambitious Plan To Beat Google's Infrastructure

semianalysis.com · 7,084 words · saved by 4 readers

Buildouts of AI infrastructure are insatiable due to the continued improvements from fueling the scaling laws. The leading frontier AI model training clusters have scaled to 100,000 GPUs this year, with 300,000+ GPUs clusters in the works for 2025. Given many physical constraints including construction timelines, permitting, regulations, and power availability, the traditional method of synchronous training of a large model at a single datacenter site are reaching a breaking point. Google, OpenAI, and Anthropic are already executing plans to expand their large model training from one site to multiple datacenter campuses. Google owns the most advanced computing systems in the world today and has pioneered the large-scale use of many critical technologies that are only just now being adopted by others such as their rack-scale liquid cooled architectures and multi-datacenter training. Gemini 1 Ultra was trained across multiple datacenters. Despite having more FLOPS available to them, thei

Multi-Datacenter Training: OpenAI's Ambitious Plan To Beat Google's Infrastructure Gigawatt Clusters, Telecom Networking, Long Haul Fiber, Hierarchical & Asynchronous SGD, Distributed Infrastructure Winners Dylan Patel , Daniel Nishball , and Jeremie Eliahou Ontiveros Sep 04, 2024 ∙ Paid 173 17 3 Share Buildouts of AI infrastructure are insatiable due to the continued improvements from fueling the scaling laws. The leading frontier AI model training clusters have scaled to 100,000 GPUs this year , with 300,000+ GPUs clusters in the works for 2025. Given many physical constraints including cons

Explore this link on the map →

saved by

related reading