flâneur

Benchmarking Coding Agents on Databricks’ Multi-Million Line Codebase | Databricks Blog

databricks.com · 1,841 words · saved by 2 readers

Databricks shares results from its internal coding benchmark, evaluating coding agents on a multi-million line codebase to optimize engineering cost and performance.

At Databricks, the way we build software is changing quickly as we aggressively adopt AI for engineering. The landscape of models and harnesses for code authoring has rapidly expanded in the last year, giving developers more choices than ever. With more options, it has become increasingly important to understand which coding agents offer the best performance on real-world coding tasks as well as understanding how task-performance varies with price. This article shares the results and methodology of the internal coding benchmark we built at Databricks, which evaluates tools on actual coding…

saved by

related reading