SWE-Bench Technical Report
At Honeycomb, we’re really excited about the future of high-quality automation in software engineering pipelines. Our mission is to build an expressive engine capable of implementing agents that can automate any step of the development stack, from code reviews and knowledge retrieval, to bug fixes and 0-to-1 development. In this report, we detail our approach to push the state-of-the-art in bug fixing through the SWE-Bench dataset, achieving 22.06% on the full dataset and 40.38% on the verified dataset. SWE-bench is a comprehensive evaluation framework designed to test language models' ability to solve real-world software engineering problems. The benchmark comprises 2294 engineering problems drawn from GitHub issues and pull requests across several open-source Python repositories. SWE-bench offers three distinct datasets: 1. Full Test Dataset (2294 instances): This complete set of problems represents a wide range of software engineering tasks, from simple bug fixes to complex feature
At Honeycomb, we’re really excited about the future of high-quality automation in software engineering pipelines. Our mission is to build an expressive engine capable of implementing agents that can automate any step of the development stack, from code reviews and knowledge retrieval, to bug fixes and 0-to-1 development. In this report, we detail our approach to push the state-of-the-art in bug fixing through the SWE-Bench dataset, achieving 22.06% on the full dataset and 40.38% on the verified dataset. SWE-bench is a comprehensive evaluation framework designed to test language models' ability
Explore this link on the map →