[2510.04371] Speculative Actions: A Lossless Framework for Faster Agentic Systems
Abstract:Despite growing interest in AI agents across industry and academia, their execution in an environment is often slow, hampering training, evaluation, and deployment. For example, a game of chess between two state-of-the-art agents may take hours. A critical bottleneck is that agent behavior unfolds sequentially: each action requires an API call, and these calls can be time-consuming. Inspired by speculative execution in microprocessors and speculative decoding in LLM inference, we propose speculative actions, a lossless framework for general agentic systems that predicts likely actions using faster models, enabling multiple steps to be executed in parallel. We evaluate this framework across three agentic environments: gaming, e-commerce, web search, and a "lossy" extension for an operating systems environment. In all cases, speculative actions achieve substantial accuracy in next-action prediction (up to 55%), translating into significant reductions in end-to-end latency. Moreover, performance can be further improved through stronger guessing models, top-K action prediction, multi-step speculation, and uncertainty-aware optimization, opening a promising path toward deploying low-latency agentic systems in the real world.
# link_2fp9b00kf4i.pdf ## Metadata - PDFFormatVersion=1.7 - IsLinearized=false - IsAcroFormPresent=false - IsXFAPresent=false - IsCollectionPresent=false - IsSignaturesPresent=false - Author=Naimeng Ye; Arnav Ahuja; Georgios Liargkovas; Yunan Lu; Kostis Kaffes; Tianyi Peng - Creator=arXiv GenPDF (tex2pdf:a6404ea) - Custom.DOI=https://doi.org/10.48550/arXiv.2510.04371 - Custom.License=http://creativecommons.org/licenses/by/4.0/ - Custom.PTEX.Fullbanner=This is pdfTeX, Version 3.141592653-2.6-1.40.28 (TeX Live 2025) kpathsea version 6.4.1 - Custom.arXivID=https://arxiv.org/abs/2510.04371v2 - Pro
Explore this link on the map →saved by
related reading
- [2410.00079] Interactive Speculative Planning: Enhance Agent Efficiency through Co-design of System and User Interfacearxiv.org
- FASTERinnovator-zero.github.io
- Speculative Decoding - philkravphilkrav.com
- AI 2027ai-2027.com
- Fast Inference from Transformers via Speculative Decodingarxiv.org
- We spent 2 hours working in the future - METRmetr.org
- Speculative Decoding - Deep Dive — ROCm Blogsrocm.blogs.amd.com
- Best practices to accelerate inference for large-scale production workloadstogether.ai
- Speculative decodingaarnphm.xyz
- Why Tool AIs Want to Be Agent AIs · Gwern.netgwern.net
- LLM Powered Autonomous Agents | Lil'Loglilianweng.github.io
- Building Effective AI Agents \ Anthropicanthropic.com