Low Latency and Model Training at Modal | Rhea Malik
rhea24.github.io · 2,824 words · saved by 3 readers
Reflections on building low-latency systems and agent-driven training at Modal
There's a very specific feeling that comes with shipping code that touches almost every single request in the product. You flip the flag and then just hold your breath. For the next few minutes, every request the product handles is running through code that I shipped. I sit there refreshing the latency dashboard every few seconds, half expecting to watch the line jump. It doesn't move. The graph looks exactly like it did an hour before, which was the best possible result, yet somehow the least satisfying one. This summer I interned at Modal, a serverless cloud platform for running…
saved by
related reading
- LLM Engineer's Almanac - Advisormodal.com
- Composer2.pdfcursor.com
- Modal: High-performance AI infrastructuremodal.com
- Prime Agent: A Self-Improving RLM Harnessarxiv.org
- Notes on the Software Factorybenedict.dev
- Reinforcement learning is an infrastructure problemmodal.com
- Introductionmodal.com
- Modal's serverless Servers | Modal Blogmodal.com
- Report: Modal Business Breakdown & Founding Story | Contrary Researchresearch.contrary.com
- RL Environments and RL for Science: Data Foundries and Multi-Agent Architecturesnewsletter.semianalysis.com
- 2025: The year in LLMssimonwillison.net
- laguna-m1-xs2-technical-report.pdfpoolside.ai