Continually improving our agent harness · Cursor
Models need a harness to become fully useful, and improving that harness means continuously iterating on context, evaluation, and model-specific tuning.
Blog / research We approach building the Cursor agent harness the way we'd approach any ambitious software product. Much of the work is vision-driven, where we start with an opinion about what the ideal agent experience should look like. From there, we form hypotheses about how to get closer to that vision, run experiments to test them, and iterate using quantitative and qualitative signals from evals and real usage. That process depends on having the right online and offline instrumentation, so we can tell when a change actually makes the harness better. When we get early access to new models
Explore this link on the map →related reading
- Effective harnesses for long-running agents \ Anthropicanthropic.com
- Composer2.pdfcursor.com
- Demystifying evals for AI agents \ Anthropicanthropic.com
- Lecture 01. Strong Models Don't Mean Reliable Execution | Learn Harness Engineeringwalkinglabs.github.io
- Harness engineering for coding agent usersmartinfowler.com
- Dynamic context discovery · Cursorcursor.com
- Building reliable AI agents · parth sareenparthsareen.com
- Best practices for coding with agents · Cursorcursor.com
- A Guide to Claude Code 2.0 and getting better at using coding agents – sankalp's blogsankalp.bearblog.dev
- [2604.25850] Agentic Harness Engineering: Observability-Driven Automatic Evolution of Coding-Agent Harnessesarxiv.org
- How to Harness Coding Agents with the Right Infrastructure | Blogalexlavaee.me
- Effective context engineering for AI agents \ Anthropicanthropic.com