flâneur — a map of the web's best reading

Raising the bar on SWE-bench Verified with Claude 3.5 Sonnet \ Anthropic

anthropic.com · 2,848 words · saved by 1 readers

A post for developers about the new Claude 3.5 Sonnet and the SWE-bench eval

Our latest model, the upgraded Claude 3.5 Sonnet , achieved 49% on SWE-bench Verified, a software engineering evaluation, beating the previous state-of-the-art model's 45%. This post explains the "agent" we built around the model, and is intended to help developers get the best possible performance out of Claude 3.5 Sonnet. SWE-bench is an AI evaluation benchmark that assesses a model's ability to complete real-world software engineering tasks. Specifically, it tests how the model can resolve GitHub issues from popular open-source Python repositories. For each task in the benchmark, the AI mod

Explore this link on the map →

saved by

related reading