How to achieve truly serverless GPUs
modal.com · 5,054 words · saved by 1 readers
A deep dive on Modal's deep tech for fast boots.
All posts Back Engineering May 12, 2026 • 20 minute read How we achieved truly serverless GPUs Charles Frye @charles_irl Member of Technical Staff Jonathan Belotti @jonobelotti_IO Member of Technical Staff Erik Bernhardsson @bernhardsson CEO and Founder Akshat Bubna @akshat_b CTO and Founder We are in the age of inference. Billion- to trillion-parameter neural networks are run on specialized accelerators at quadrillions of operations per second to generate media , author software , and fold proteins at massive scale. Inference workloads are more variable and less predictable than the training
saved by
related reading
- Modal: High-performance AI infrastructuremodal.com
- LLM Engineer's Almanac - Workloads | Modalmodal.com
- The Best GPUs for Deep Learning in 2023 — An In-depth Analysistimdettmers.com
- Modal's serverless Servers | Modal Blogmodal.com
- Together AI | The AI Native Cloudtogether.ai
- LLM Inference Performance Engineering: Best Practices | Databricks Blogdatabricks.com
- Look Ma, No Bubbles! Designing a Low-Latency Megakernel for Llama-1B · Hazy Researchhazyresearch.stanford.edu
- ⭐️ Fast LLM Inference From Scratchandrewkchan.dev
- Introductionmodal.com
- The Short Case for Nvidia Stock | YouTube Transcript Optimizeryoutubetranscriptoptimizer.com
- Real-time LLM Inference on Standard Datacenter GPUs (3,000 tokens/s per request)blog.kog.ai
- Low Latency and Model Training at Modalrhea24.github.io