Raising the bar on SWE-bench Verified with Claude 3.5 Sonnet \ Anthropic
anthropic.com · 2,848 words · saved by 2 readers
A post for developers about the new Claude 3.5 Sonnet and the SWE-bench eval
Our latest model, the upgraded Claude 3.5 Sonnet , achieved 49% on SWE-bench Verified, a software engineering evaluation, beating the previous state-of-the-art model's 45%. This post explains the "agent" we built around the model, and is intended to help developers get the best possible performance out of Claude 3.5 Sonnet. SWE-bench is an AI evaluation benchmark that assesses a model's ability to complete real-world software engineering tasks. Specifically, it tests how the model can resolve GitHub issues from popular open-source Python repositories. For each task in the benchmark, the AI mod
saved by
related reading
- A Guide to Claude Code 2.0 and getting better at using coding agents – sankalp's blogsankalp.bearblog.dev
- Cookbookcookbook.openai.com
- Composer2.pdfcursor.com
- Best practices for Claude Code - Claude Code Docsanthropic.com
- PostTrainBenchposttrainbench.com
- FrontierSWEfrontierswe.com
- Senior SWE-Benchsenior-swe-bench.snorkel.ai
- Claude Code Cheat Sheetcc.storyfox.cz
- 2502.18449arxiv.org
- Claude Code Opus 4.8 Performance Tracker | Marginlabmarginlab.ai
- FrontierSWEfrontierswe.com
- [2410.06992] SWE-Bench+: Enhanced Coding Benchmark for LLMsar5iv.labs.arxiv.org