flâneur

Manifesto | Sail Research

sailresearch.com · 475 words · saved by 1 readers

Inference + sandboxes for long-horizon agents

The inference behind every AI workload makes a trade-off between latency and throughput. In the first iteration of generative AI, systems optimized for low latency return tokens and output as quickly as possible to a user waiting on the other end. But speed comes at a cost. A large and growing share of AI work isn’t waiting on a human at all. Asynchronous use cases—like deep research, code review, security review, evals, and embeddings—require agentic pipelines that spend hours running in the background, without humans in the loop. In this paradigm, shaving milliseconds off a single…

saved by

related reading