flâneur — a map of the web's best reading

Network and Storage Benchmarks for LLM Training on the Cloud | Henry Zhu

maknee.github.io · 1,997 words · saved by 1 readers

Personal website for some random tidbits I work on

AI usage has become universal. Teams everywhere are building RAG, generating embeddings, and training increasingly sophisticated agents. Most distributed LLM training guides focus on model architecture and hyperparameters while ignoring a critical bottleneck: infrastructure configuration. Network and storage choices often determine whether training takes hours or days. I ran benchmarks finetuning Gemma 3 12B and GPT-OSS-120B with different storage and network configurations using SkyPilot for infra and Nebius for GPUs. The results reveal that InfiniBand networking provides 10x faster training

Explore this link on the map →

related reading