✳flâneur — a map of the web's best reading
Raising the bar on SWE-bench Verified with Claude 3.5 Sonnet \ Anthropic
anthropic.com · 2,848 words · saved by 1 readers
A post for developers about the new Claude 3.5 Sonnet and the SWE-bench eval
Our latest model, the upgraded Claude 3.5 Sonnet , achieved 49% on SWE-bench Verified, a software engineering evaluation, beating the previous state-of-the-art model's 45%. This post explains the "agent" we built around the model, and is intended to help developers get the best possible performance out of Claude 3.5 Sonnet. SWE-bench is an AI evaluation benchmark that assesses a model's ability to complete real-world software engineering tasks. Specifically, it tests how the model can resolve GitHub issues from popular open-source Python repositories. For each task in the benchmark, the AI mod
Explore this link on the map →saved by
related reading
- AI Benchmark Leaderboards & Model Evals | BenchmarkListbenchmarklist.com
- PostTrainBenchposttrainbench.com
- Claude Sonnet 4.5 System Cardassets.anthropic.com
- A Guide to Claude Code 2.0 and getting better at using coding agents – sankalp's blogsankalp.bearblog.dev
- Cookbookcookbook.openai.com
- Composer2.pdfcursor.com
- Best practices for Claude Code - Claude Code Docsanthropic.com
- Claude Code Cheat Sheetcc.storyfox.cz
- 2502.18449arxiv.org
- Claude Code Opus 4.8 Performance Tracker | Marginlabmarginlab.ai
- How I Use Claude Code | Philipp Spiessspiess.dev
- [2410.06992] SWE-Bench+: Enhanced Coding Benchmark for LLMsar5iv.labs.arxiv.org