MirrorCode: Evidence AI can already do some weeks-long coding tasks | Epoch AI
epoch.ai · 11,210 words · saved by 1 readers
Early results from MirrorCode benchmark with METR: AI agents can complete weeks-long coding tasks, including reimplementing a 16,000-line codebase.
MirrorCode: Evidence AI can already do some weeks-long coding tasks | Epoch AI Introduction We present early results from MirrorCode , a benchmark (co-developed with METR ) of long-horizon coding tasks derived from real software applications. We find that AI models can autonomously reimplement complex existing software without access to the original program’s source code, provided there is a detailed, checkable specification. For example, Claude Opus 4.6 successfully reimplemented gotree — a bioinformatics toolkit with ~16,000 lines of Go and 40+ commands. We guess this same task would take a
saved by
related reading
- Frontier Coding Agents Can Now Implement an AlphaZero Self-Play Machine Learning Pipeline For Connect Four That Performs Comparably to an External Solver — LessWronglesswrong.com
- How to Harness Coding Agents with the Right Infrastructure | Blogalexlavaee.me
- Composer2.pdfcursor.com
- A Guide to Claude Code 2.0 and getting better at using coding agents – sankalp's blogsankalp.bearblog.dev
- Shipping at Inference-Speed | Peter Steinbergersteipete.me
- Best practices for Claude Code - Claude Code Docsanthropic.com
- Effective harnesses for long-running agents \ Anthropicanthropic.com
- What it Takes for Coding Agents to Complete Large Software Tasksfactory.ai
- Benchmarking Coding Agents on Databricks’ Multi-Million Line Codebasedatabricks.com
- Agent Leaderboards · Which tools coding agents choose · Armaturearmature.tech
- 2025: The year in LLMssimonwillison.net
- How coding agents read your code (and how to write for them)modem.dev