✳flâneur — a map of the web's best reading
MirrorCode: Evidence AI can already do some weeks-long coding tasks | Epoch AI
epoch.ai · 11,210 words · saved by 1 readers
Early results from MirrorCode benchmark with METR: AI agents can complete weeks-long coding tasks, including reimplementing a 16,000-line codebase.
MirrorCode: Evidence AI can already do some weeks-long coding tasks | Epoch AI Introduction We present early results from MirrorCode , a benchmark (co-developed with METR ) of long-horizon coding tasks derived from real software applications. We find that AI models can autonomously reimplement complex existing software without access to the original program’s source code, provided there is a detailed, checkable specification. For example, Claude Opus 4.6 successfully reimplemented gotree — a bioinformatics toolkit with ~16,000 lines of Go and 40+ commands. We guess this same task would take a
Explore this link on the map →saved by
related reading
- Frontier Coding Agents Can Now Implement an AlphaZero Self-Play Machine Learning Pipeline For Connect Four That Performs Comparably to an External Solver — LessWronglesswrong.com
- How to Harness Coding Agents with the Right Infrastructure | Blogalexlavaee.me
- Composer2.pdfcursor.com
- A Guide to Claude Code 2.0 and getting better at using coding agents – sankalp's blogsankalp.bearblog.dev
- Shipping at Inference-Speed | Peter Steinbergersteipete.me
- Best practices for Claude Code - Claude Code Docsanthropic.com
- Effective harnesses for long-running agents \ Anthropicanthropic.com
- 2025: The year in LLMssimonwillison.net
- How I'm using coding agents in September, 2025 — Massively Parallel Procrastinationblog.fsck.com
- ProgramBenchprogrambench.com
- Introducing FrontierCode | Cognitioncognition.ai
- Coding Models Are Doing Too Much | whnrehiew.github.io