✳flâneur — a map of the web's best reading
How to achieve truly serverless GPUs
modal.com · 5,054 words · saved by 1 readers
A deep dive on Modal's deep tech for fast boots.
All posts Back Engineering May 12, 2026 • 20 minute read How we achieved truly serverless GPUs Charles Frye @charles_irl Member of Technical Staff Jonathan Belotti @jonobelotti_IO Member of Technical Staff Erik Bernhardsson @bernhardsson CEO and Founder Akshat Bubna @akshat_b CTO and Founder We are in the age of inference. Billion- to trillion-parameter neural networks are run on specialized accelerators at quadrillions of operations per second to generate media , author software , and fold proteins at massive scale. Inference workloads are more variable and less predictable than the training
Explore this link on the map →saved by
related reading
- LLM Engineer's Almanac - Workloads | Modalmodal.com
- The Best GPUs for Deep Learning in 2023 — An In-depth Analysistimdettmers.com
- Transformer Inference Arithmetic | kipply's blogkipp.ly
- LLM Inference Performance Engineering: Best Practices | Databricks Blogdatabricks.com
- Real-time LLM Inference on Standard Datacenter GPUs (3,000 tokens/s per request)blog.kog.ai
- ⭐️ Fast LLM Inference From Scratchandrewkchan.dev
- Look Ma, No Bubbles! Designing a Low-Latency Megakernel for Llama-1B · Hazy Researchhazyresearch.stanford.edu
- The Short Case for Nvidia Stock | YouTube Transcript Optimizeryoutubetranscriptoptimizer.com
- Best practices to accelerate inference for large-scale production workloadstogether.ai
- Report: Modal Business Breakdown & Founding Story | Contrary Researchresearch.contrary.com
- GPU Instances and Serverless Inference — Verda (formerly DataCrunch)verda.com
- How to Think About GPUs | How To Scale Your Modeljax-ml.github.io