Training an Agentic Router for Optimal Cost-Performance on SWE Tasks | Applied Compute
On most enterprise tasks, model quality is not a scalar. One model is better at long-horizon repository exploration. Another is better at small, surgical patches. A third has stronger general reasoning but higher latency or cost. The right model depends on the task, the environment, and the failure mode that matters. This is especially apparent in agentic software engineering, where every task unfolds as a trajectory. The model reads files, searches the repository, forms a hypothesis, edits code, runs tests, and decides when to stop. The best model for one issue may be the wrong model for the next one. Routing each task to the model best suited to solve it captures the strengths of specialized models and routes around their weaknesses, instead of settling for one model's compromises across the board. But this only works with a router that picks the model most likely to solve a given task, at the lowest cost, using the local evidence available before the rollout begins. Many of our work
On most enterprise tasks, model quality is not a scalar. One model is better at long-horizon repository exploration. Another is better at small, surgical patches. A third has stronger general reasoning but higher latency or cost. The right model depends on the task, the environment, and the failure mode that matters. This is especially apparent in agentic software engineering, where every task unfolds as a trajectory. The model reads files, searches the repository, forms a hypothesis, edits code, runs tests, and decides when to stop. The best model for one issue may be the wrong model for the
Explore this link on the map →related reading
- Composer2.pdfcursor.com
- AI Model & API Providers Analysis | Artificial Analysisartificialanalysis.ai
- RouteLLM: An Open-Source Framework for Cost-Effective LLM Routing - LMSYS Orglmsys.org
- Agent swarms and the new model economics · Cursorcursor.com
- Introducing SWE-grep and SWE-grep-mini: RL for Multi-Turn, Fast Context Retrieval | Cognitioncognition.ai
- Building Effective AI Agents \ Anthropicanthropic.com
- Language Models can Solve Computer Tasksarxiv.org
- Building Effective AI Agents \ Anthropicanthropic.com
- Quantifying infrastructure noise in agentic coding evals \ Anthropicanthropic.com
- Building reliable AI agents · parth sareenparthsareen.com
- PostTrainBenchposttrainbench.com
- Arjun Virkarjunvirk.com