How we compare model quality in Cursor · Cursor
cursor.com · 1,044 words · saved by 2 readers
We use a hybrid online-offline eval process to keep our understanding of model quality aligned with what developers actually do.
Note: CursorBench is continually updated as agent capabilities evolve. The current production version is CursorBench 3.1; see that page for the latest leaderboard. Developers are asking coding agents to take on longer, more complex tasks that span multiple files, tools, and steps. As these requests grow in scope, the evals that measure agent performance need to evolve with them. At Cursor, we use a hybrid online-offline eval process to keep our understanding of model quality aligned with what developers actually do. The offline part uses CursorBench, our internal eval suite based on real…
saved by
related reading
- Demystifying evals for AI agents \ Anthropicanthropic.com
- Benchmarking Coding Agents on Databricks’ Multi-Million Line Codebasedatabricks.com
- Composer2.pdfcursor.com
- Coding vs thinking — Paradigm 3paradigm3.org
- FrontierSWEfrontierswe.com
- Continually improving our agent harness · Cursorcursor.com
- Introducing FrontierCode | Cognitioncognition.com
- Cursor · CursorBenchcursor.com
- Introducing FrontierCode | Cognitioncognition.ai
- WebCode: Search Evals for Coding Agentsexa.ai
- Senior SWE-Benchsenior-swe-bench.snorkel.ai
- Cursor: AI coding agentcursor.sh