HPC vs AI
System architecture There are a few differences between designing a supercomputer for AI and designing a supercomputer for traditional modeling and simulation.
System architecture There are a few differences between designing a supercomputer for AI and designing a supercomputer for traditional modeling and simulation. I once gave a presentation internally at Microsoft that contained this table: The biggest differences amount to: AI benefits from multi-plane fat trees. They make collectives faster since more nodes can talk to each other without hopping through switches, but they require more nodes and switches to be connected together in a small space. In practice, this requires using expensive optics instead of DAC cables. AI uses a lot of node-local
Explore this link on the map →saved by
related reading
- How To Scale Your Modeljax-ml.github.io
- Multi-Datacenter Training: OpenAI's Ambitious Plan To Beat Google's Infrastructuresemianalysis.com
- Components of an Open Source AI Compute Tech Stackanyscale.com
- My picture of the present in AI — LessWronglesswrong.com
- Reiner Pope – The math behind how LLMs are trained and serveddwarkesh.com
- The Short Case for Nvidia Stock | YouTube Transcript Optimizeryoutubetranscriptoptimizer.com
- How to Think About GPUs | How To Scale Your Modeljax-ml.github.io
- A Hitchhiker’s Guide to ML Training Infrastructure | CMU Software Engineering Institutesei.cmu.edu
- Training great LLMs entirely from ground up in the wilderness as a startup - Yi Tayyitay.net
- An Interview with MatX CEO Reiner Pope About LLM Chipschipstrat.com
- Can AI scaling continue through 2030? | Epoch AIepoch.ai
- Accelerate AI & Machine Learning Workflows | NVIDIA Run:airun.ai