SkyNomad: On Using Multi-Region Spot Instances to Minimize AI Batch Job Cost
This is experimental HTML to improve accessibility. We invite you to report rendering errors. Use Alt+Y to toggle on accessible reporting links and Alt+Shift+Y to toggle off. Learn more about this project and help improve conversions. spacing=nonfrench AI batch jobs such as model training, inference pipelines, and data analytics require substantial GPU resources and often need to finish before a deadline. Spot instances offer 3 – 10 × lower cost than on-demand instances, but their unpredictable availability makes meeting deadlines difficult. Existing systems either rely solely on spot instances and risk deadline violations, or operate in simplified single-region settings. These approaches overlook substantial spatial and temporal heterogeneity in spot availability, lifetimes, and prices. We show that exploiting such heterogeneity to access more spot capacity is the key to reduce the job execution cost. We present SkyNomad , a multi-region scheduling system that maximizes spot usage
spacing=nonfrench SkyNomad: On Using Multi-Region Spot Instances to Minimize AI Batch Job Cost Zhifei Li ∗† , Tian Xia ∗† , Ziming Mao † , Zihan Zhou ¶ , Ethan J. Jackson † , Jamison Kerney † , Zhanghao Wu † , Pratik Mishra § , Yi Xu † , Yifan Qiao † , Scott Shenker †,⋄ , Ion Stoica † † UC Berkeley ¶ Shanghai Jiao Tong University § AMD ⋄ ICSI Abstract AI batch jobs such as model training, inference pipelines, and data analytics require substantial GPU resources and often need to finish before a deadline. Spot instances offer 3 – 10 × 3\text{--}10\times lower cost than on-demand instances, but
Explore this link on the map →saved by
related reading
- AWS Batch :: EC2 Spot Workshopsec2spotworkshops.com
- Getting $1M cloud credits for AI startups — and using them wisely | SkyPilot Blogblog.skypilot.co
- My picture of the present in AI — LessWronglesswrong.com
- AI Model & API Providers Analysis | Artificial Analysisartificialanalysis.ai
- Designing Compute Markets | Kavish Gargkavishgarg.com
- Multi-Datacenter Training: OpenAI's Ambitious Plan To Beat Google's Infrastructuresemianalysis.com
- To Boldly Go: The Case for Space Datacentersnewsletter.semianalysis.com
- Fast, scalable, clean, and cheap enoughoffgridai.us
- Navigating the High Cost of AI Compute | Andreessen Horowitza16z.com
- Reiner Pope – The math behind how LLMs are trained and serveddwarkesh.com
- Cheap RL tasks will waste compute | Mechanize, Inc.mechanize.work
- LLM Engineer's Almanac - Workloads | Modalmodal.com