[2410.06992] SWE-Bench+: Enhanced Coding Benchmark for LLMs
Large Language Models (LLMs) in Software Engineering (SE) can offer assistance for coding. To facilitate a rigorous evaluation of LLMs in practical coding contexts, Carlos et al. introduced the SWE-bench dataset, which…
SWE-Bench+: Enhanced Coding Benchmark for LLMs Reem Aleithan, Haoran Xue, Mohammad Mahdi Mohajer, Elijah Nnorom, Gias Uddin, Song Wang Lassonde School of Engineering York University {reem1100,hrx00,mmm98,ennorom,guddin,wangsong}@yorku.ca Abstract Large Language Models (LLMs) in Software Engineering (SE) can offer assistance for coding. To facilitate a rigorous evaluation of LLMs in practical coding contexts, Carlos et al. introduced the SWE-bench dataset, which comprises 2,294 real-world GitHub issues and their corresponding pull requests, collected from 12 widely used Python repositories. Sev
Explore this link on the map →related reading
- 2502.18449arxiv.org
- Claude SWE-Bench Performance \ Anthropicanthropic.com
- Composer2.pdfcursor.com
- Introducing FrontierCode | Cognitioncognition.ai
- 2025: The year in LLMssimonwillison.net
- crawshaw - 2025-01-06crawshaw.io
- Coding Models Are Doing Too Much | whnrehiew.github.io
- LiCoEval: Evaluating LLMs on License Compliance in Code Generationarxiv.org
- Here’s how I use LLMs to help me write codesimonwillison.net
- [2501.01257] CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratingsarxiv.org
- Introducing FrontierCode | Cognitioncognition.com
- Senior SWE-Benchsenior-swe-bench.snorkel.ai