Modal's serverless Servers
modal.com · 2,304 words · saved by 1 readers
A deep dive inside our new ultra-low-latency primitive.
Modal makes it easy to run high-performance code in the cloud: Python functions, agent runtimes, notebooks, batch jobs, and more. Now, you can also run ultra-low-latency Servers on Modal for HTTP, WebSocket, and gRPC traffic. Servers are designed for applications where every millisecond counts, like LLM inference for interactive agents. Servers give you a regionalized, autoscaling pool of HTTP server replicas behind Modal’s routing layer, with the deployment ergonomics, fast feedback loops, and autoscaling we consider table stakes (for humans and for agents). This might sound familiar:…
saved by
related reading
- Modal is a computer | Modal Blogmodal.com
- Modal: High-performance AI infrastructuremodal.com
- Low Latency and Model Training at Modalrhea24.github.io
- Introductionmodal.com
- Lambda on hard mode: Inside Modal's web infrastructuremodal.com
- How we achieved truly serverless GPUsmodal.com
- What I have been working on: Modal · Erik Bernhardssonerikbern.com
- Together AI | The AI Native Cloudtogether.ai
- LLM Engineer's Almanac - Workloads | Modalmodal.com
- Dynamic batching | Modal Docsmodal.com
- Real-time inference for robots at Physical Intelligence | Modal Blogmodal.com
- Modal (@modal) on Xx.com