[2410.06992] SWE-Bench+: Enhanced Coding Benchmark for LLMs
Large Language Models (LLMs) in Software Engineering (SE) can offer assistance for coding. To facilitate a rigorous evaluation of LLMs in practical coding contexts, Carlos et al. introduced the SWE-bench dataset, which…
SWE-Bench+: Enhanced Coding Benchmark for LLMs Reem Aleithan, Haoran Xue, Mohammad Mahdi Mohajer, Elijah Nnorom, Gias Uddin, Song Wang Lassonde School of Engineering York University {reem1100,hrx00,mmm98,ennorom,guddin,wangsong}@yorku.ca Abstract Large Language Models (LLMs) in Software Engineering (SE) can offer assistance for coding. To facilitate a rigorous evaluation of LLMs in practical coding contexts, Carlos et al. introduced the SWE-bench dataset, which comprises 2,294 real-world GitHub issues and their corresponding pull requests, collected from 12 widely used Python repositories. Sev
related reading
- 2502.18449arxiv.org
- Claude SWE-Bench Performance \ Anthropicanthropic.com
- Senior SWE-Benchsenior-swe-bench.snorkel.ai
- Composer2.pdfcursor.com
- SWE-chat: Coding Agent Interactions From Real Users in the Wildarxiv.org
- crawshaw - 2025-01-06crawshaw.io
- Introducing FrontierCode | Cognitioncognition.ai
- Introducing FrontierCode | Cognitioncognition.com
- Lean Software Scaling Laws · Gwern.netgwern.net
- 2025: The year in LLMssimonwillison.net
- FrontierSWEfrontierswe.com
- BigO(Bench) -- Can LLMs Generate Code with Controlled Time and Space Complexity?alphaxiv.org