flâneur

How we compare model quality in Cursor · Cursor

cursor.com · 1,044 words · saved by 2 readers

We use a hybrid online-offline eval process to keep our understanding of model quality aligned with what developers actually do.

Note: CursorBench is continually updated as agent capabilities evolve. The current production version is CursorBench 3.1; see that page for the latest leaderboard. Developers are asking coding agents to take on longer, more complex tasks that span multiple files, tools, and steps. As these requests grow in scope, the evals that measure agent performance need to evolve with them. At Cursor, we use a hybrid online-offline eval process to keep our understanding of model quality aligned with what developers actually do. The offline part uses CursorBench, our internal eval suite based on real…

saved by

related reading