flâneur — a map of the web's best reading

[2410.06992] SWE-Bench+: Enhanced Coding Benchmark for LLMs

ar5iv.labs.arxiv.org · 7,017 words · saved by 1 readers

Large Language Models (LLMs) in Software Engineering (SE) can offer assistance for coding. To facilitate a rigorous evaluation of LLMs in practical coding contexts, Carlos et al. introduced the SWE-bench dataset, which…

SWE-Bench+: Enhanced Coding Benchmark for LLMs Reem Aleithan, Haoran Xue, Mohammad Mahdi Mohajer, Elijah Nnorom, Gias Uddin, Song Wang Lassonde School of Engineering York University {reem1100,hrx00,mmm98,ennorom,guddin,wangsong}@yorku.ca Abstract Large Language Models (LLMs) in Software Engineering (SE) can offer assistance for coding. To facilitate a rigorous evaluation of LLMs in practical coding contexts, Carlos et al. introduced the SWE-bench dataset, which comprises 2,294 real-world GitHub issues and their corresponding pull requests, collected from 12 widely used Python repositories. Sev

Explore this link on the map →

related reading