Multi-Datacenter Training: OpenAI's Ambitious Plan To Beat Google's Infrastructure
Buildouts of AI infrastructure are insatiable due to the continued improvements from fueling the scaling laws. The leading frontier AI model training clusters have scaled to 100,000 GPUs this year, with 300,000+ GPUs clusters in the works for 2025. Given many physical constraints including construction timelines, permitting, regulations, and power availability, the traditional method of synchronous training of a large model at a single datacenter site are reaching a breaking point. Google, OpenAI, and Anthropic are already executing plans to expand their large model training from one site to multiple datacenter campuses. Google owns the most advanced computing systems in the world today and has pioneered the large-scale use of many critical technologies that are only just now being adopted by others such as their rack-scale liquid cooled architectures and multi-datacenter training. Gemini 1 Ultra was trained across multiple datacenters. Despite having more FLOPS available to them, thei
Multi-Datacenter Training: OpenAI's Ambitious Plan To Beat Google's Infrastructure Gigawatt Clusters, Telecom Networking, Long Haul Fiber, Hierarchical & Asynchronous SGD, Distributed Infrastructure Winners Dylan Patel , Daniel Nishball , and Jeremie Eliahou Ontiveros Sep 04, 2024 ∙ Paid 173 17 3 Share Buildouts of AI infrastructure are insatiable due to the continued improvements from fueling the scaling laws. The leading frontier AI model training clusters have scaled to 100,000 GPUs this year , with 300,000+ GPUs clusters in the works for 2025. Given many physical constraints including cons
Explore this link on the map →saved by
related reading
- Can AI scaling continue through 2030? | Epoch AIepoch.ai
- How To Scale Your Modeljax-ml.github.io
- How to Build the Future of AI in the United States | IFPifp.org
- Training great LLMs entirely from ground up in the wilderness as a startup - Yi Tayyitay.net
- The Short Case for Nvidia Stock | YouTube Transcript Optimizeryoutubetranscriptoptimizer.com
- IIIa. Racing to the Trillion-Dollar Cluster - SITUATIONAL AWARENESSsituational-awareness.ai
- Can AI scaling continue through 2030? | Epoch AIepochai.org
- A Hitchhiker’s Guide to ML Training Infrastructure | CMU Software Engineering Institutesei.cmu.edu
- My picture of the present in AI — LessWronglesswrong.com
- From bare metal to a 70B model: infrastructure set-up and scripts - Imbueimbue.com
- How to Think About GPUs | How To Scale Your Modeljax-ml.github.io
- The Best GPUs for Deep Learning in 2023 — An In-depth Analysistimdettmers.com