WebCode: Search Evals for Coding Agents | Exa Blog
exa.ai · 1,728 words · saved by 1 readers
How do you measure retrieval quality for the increasingly complex ecosystem coding agents consume from?
#State of code search Today, we're open-sourcing a set of coding evaluations, WebCode, that we built to evaluate web search for coding agents. At Exa, we power search for most of the largest coding agent companies, and we have observed a surge in code search queries over the past year, with a particularly large jump at the end of 2025. Figure 1: Code search queries on Exa This growth pushed us to focus on code search, where precision is especially important: agents build on retrieved context across many steps, so stale or noisy search results can poison or even derail the reasoning…
saved by
related reading
- Demystifying evals for AI agents \ Anthropicanthropic.com
- Introducing SWE-grep and SWE-grep-mini: RL for Multi-Turn, Fast Context Retrieval | Cognitioncognition.ai
- Composer2.pdfcursor.com
- Our AI Research: How We Evaluate Semantic Search Technology | Exa Blogexa.ai
- NEEDLE: The benchmark your search engine can't memorize - Keenable.aikeenable.ai
- Use exa-code: Fast, efficient web context for coding agents | Exa Blogexa.ai
- Chroma Context-1: Training a Self-Editing Search Agent | Chromatrychroma.com
- The Verification Horizon: No Silver Bullet for Coding Agent Rewardsarxiv.org
- How we compare model quality in Cursorcursor.com
- Senior SWE-Benchsenior-swe-bench.snorkel.ai
- Patterns for Building LLM-based Systems & Productseugeneyan.com
- SWE-chat: Coding Agent Interactions From Real Users in the Wildarxiv.org