Harness design for long-running application development \ Anthropic
anthropic.com · 5,100 words · saved by 4 readers
Anthropic is an AI safety and research company that's working to build reliable, interpretable, and steerable AI systems.
Written by Prithvi Rajasekaran, a member of our Labs team. Over the past several months I’ve been working on two interconnected problems: getting Claude to produce high-quality frontend designs, and getting it to build complete applications without human intervention. This work originated with earlier efforts on our frontend design skill and long-running coding agent harness, where my colleagues and I were able to improve Claude’s performance well above baseline through prompt engineering and harness design—but both eventually hit ceilings. To break through, I sought out novel AI…
saved by
related reading
- Effective harnesses for long-running agents \ Anthropicanthropic.com
- Demystifying evals for AI agents \ Anthropicanthropic.com
- Welcome to Learn Harness Engineeringwalkinglabs.github.io
- Harness Engineering for Self-Improvement | Lil'Loglilianweng.github.io
- Building Effective AI Agents \ Anthropicanthropic.com
- A Guide to Claude Code 2.0 and getting better at using coding agents – sankalp's blogsankalp.bearblog.dev
- Trending Papers - Hugging Facepaperswithcode.com
- Using AI as a Design Engineerjakub.kr
- The Shape of AI | UX Patterns for Artificial Intelligence Designshapeof.ai
- Building Effective AI Agents \ Anthropicanthropic.com
- Scaling Managed Agents: Decoupling the brain from the hands \ Anthropicanthropic.com
- Towards self-driving codebases · Cursorcursor.com