Benchmarking Coding Agents on Databricks’ Multi-Million Line Codebase | Databricks Blog
Databricks shares results from its internal coding benchmark, evaluating coding agents on a multi-million line codebase to optimize engineering cost and performance.
At Databricks, the way we build software is changing quickly as we aggressively adopt AI for engineering. The landscape of models and harnesses for code authoring has rapidly expanded in the last year, giving developers more choices than ever. With more options, it has become increasingly important to understand which coding agents offer the best performance on real-world coding tasks as well as understanding how task-performance varies with price. This article shares the results and methodology of the internal coding benchmark we built at Databricks, which evaluates tools on actual coding…
saved by
related reading
- Composer2.pdfcursor.com
- How coding agents read your code (and how to write for them)modem.dev
- How we compare model quality in Cursorcursor.com
- AI Model & API Providers Analysis | Artificial Analysisartificialanalysis.ai
- Shipping at Inference-Speed | Peter Steinbergersteipete.me
- [2604.22750] How Do AI Agents Spend Your Money? Analyzing and Predicting Token Consumption in Agentic Coding Tasksarxiv.org
- Quantifying infrastructure noise in agentic coding evals \ Anthropicanthropic.com
- AINews | AINewsnews.smol.ai
- Introducing FrontierCode | Cognitioncognition.ai
- Introducing FrontierCode | Cognitioncognition.com
- 2025: The year in LLMssimonwillison.net
- MirrorCode: Evidence AI can already do some weeks-long coding tasks | Epoch AIepoch.ai