[2604.25850] Agentic Harness Engineering: Observability-Driven Automatic Evolution of Coding-Agent Harnesses
Abstract:Harnesses are now central to coding-agent performance, mediating how models interact with tools and execution environments. Yet harness engineering remains a manual craft, because automating it faces a heterogeneous action space across editable components, voluminous trajectories that bury actionable signal, and edits whose effect is hard to attribute. We introduce Agentic Harness Engineering (AHE), a closed loop that addresses these challenges through three matched observability pillars: (1) component observability gives every editable harness component a file-level representation so the action space is explicit and revertible; (2) experience observability distills millions of raw trajectory tokens into a layered, drill-down evidence corpus that an evolving agent can actually consume; and (3) decision observability pairs every edit with a self-declared prediction, later verified against the next round's task-level outcomes. Together, these pillars turn every edit into a falsifiable contract, so harness evolution proceeds autonomously without collapsing into trial-and-error. Empirically, ten AHE iterations lift pass@1 on Terminal-Bench 2 from 69.7% to 77.0%, surpassing the human-designed harness Codex-CLI (71.9%) and the self-evolving baselines ACE and TF-GRPO. The frozen harness transfers without re-evolution: on SWE-bench-verified it tops aggregate success at 12% fewer tokens than the seed, and on Terminal-Bench 2 it yields +5.1 to +10.1pp cross-family gains across three alternate model families, indicating the evolved components encode general engineering experience rather than benchmark-specific tuning. Ablations localize the gain to tools, middleware, and long-term memory rather than the system prompt, suggesting factual harness structure transfers while prose-level strategy does not.
[2604.25850] Agentic Harness Engineering: Observability-Driven Automatic Evolution of Coding-Agent Harnesses --> Computer Science > Computation and Language arXiv:2604.25850 (cs) [Submitted on 28 Apr 2026 ( v1 ), last revised 18 May 2026 (this version, v4)] Title: Agentic Harness Engineering: Observability-Driven Automatic Evolution of Coding-Agent Harnesses Authors: Jiahang Lin , Shichun Liu , Chengjun Pan , Lizhi Lin , Shihan Dou , Zhiheng Xi , Xuanjing Huang , Hang Yan , Zhenhua Han , Tao Gui , Yu-Gang Jiang View a PDF of the paper titled Agentic Harness Engineering: Observability-Driven Au
Explore this link on the map →saved by
related reading
- Harness engineering for coding agent usersmartinfowler.com
- How to Harness Coding Agents with the Right Infrastructure | Blogalexlavaee.me
- Lecture 01. Strong Models Don't Mean Reliable Execution | Learn Harness Engineeringwalkinglabs.github.io
- Demystifying evals for AI agents \ Anthropicanthropic.com
- Effective harnesses for long-running agents \ Anthropicanthropic.com
- [2606.09498] Self-Harness: Harnesses That Improve Themselvesarxiv.org
- Harness Engineering for Self-Improvement | Lil'Loglilianweng.github.io
- A Guide to Claude Code 2.0 and getting better at using coding agents – sankalp's blogsankalp.bearblog.dev
- sysls on X: "How To Be A World-Class Agentic Engineer" / Xx.com
- [2606.09498] Self-Harness: Harnesses That Improve Themselvesarxiv.org
- Composer2.pdfcursor.com
- Continually improving our agent harness · Cursorcursor.com