Agent Leaderboards · Which tools coding agents choose · Armature
armature.tech · 259 words · saved by 1 readers
Which solutions coding agents choose, measured with controlled experiments. Replay every run.
Every number comes from a controlled experiment. Same repositories, same prompts, real agents, judged results. Amplifying's greenfield benchmark measures open-ended recommendations. These boards also require an approved implementation and a production operating path. Community-baseline asks are therefore shown separately from deployment- and repository-constrained work; that seam is the closest like-for-like comparison. 01 A panel of repositories Synthetic repositories that look like real companies. Different languages, stacks and personas. Every run starts from the same code. 02 Real…
saved by
related reading
- Inside Coding Agentsvpromise.github.io
- Agentationagentation.dev
- A Guide to Claude Code 2.0 and getting better at using coding agents – sankalp's blogsankalp.bearblog.dev
- Best practices for Claude Code - Claude Code Docsanthropic.com
- Agentationagentation.com
- Why We Built Our Own Background Agentbuilders.ramp.com
- Replicas | Cloud Coding Agents for Engineering Teamsreplicas.dev
- Effective harnesses for long-running agents \ Anthropicanthropic.com
- How I'm using coding agents in September, 2025 — Massively Parallel Procrastinationblog.fsck.com
- GitHub - x1xhlol/system-prompts-and-models-of-ai-tools: FULL Augment Code, Claude Code, Cluely, CodeBuddy, Comet, Cursor, Devin AI, Junie, Kiro, Leap.new, Lovable, Manus, NotionAI, Orchids.app, Perplexity, Poke, Qoder, Replit, Same.dev, Trae, Traycer AI, VSCode Agent, Warp.dev, Windsurf, Xcode, Z.ai Code, Dia & v0. (And other Open Sourced) System Prompts, Internal Tools & AI Modelsgithub.com
- Designing agentic loopssimonwillison.net
- SWE-chat: Coding Agent Interactions From Real Users in the Wildarxiv.org