flâneur — a map of the web's best reading

SWE-Bench Technical Report

honeycomb.sh · saved by 1 readers

At Honeycomb, we’re really excited about the future of high-quality automation in software engineering pipelines. Our mission is to build an expressive engine capable of implementing agents that can automate any step of the development stack, from code reviews and knowledge retrieval, to bug fixes and 0-to-1 development. In this report, we detail our approach to push the state-of-the-art in bug fixing through the SWE-Bench dataset, achieving 22.06% on the full dataset and 40.38% on the verified dataset. SWE-bench is a comprehensive evaluation framework designed to test language models' ability to solve real-world software engineering problems. The benchmark comprises 2294 engineering problems drawn from GitHub issues and pull requests across several open-source Python repositories. SWE-bench offers three distinct datasets: 1. Full Test Dataset (2294 instances): This complete set of problems represents a wide range of software engineering tasks, from simple bug fixes to complex feature

At Honeycomb, we’re really excited about the future of high-quality automation in software engineering pipelines. Our mission is to build an expressive engine capable of implementing agents that can automate any step of the development stack, from code reviews and knowledge retrieval, to bug fixes and 0-to-1 development. In this report, we detail our approach to push the state-of-the-art in bug fixing through the SWE-Bench dataset, achieving 22.06% on the full dataset and 40.38% on the verified dataset. SWE-bench is a comprehensive evaluation framework designed to test language models' ability

Explore this link on the map →

saved by