What it Takes for Coding Agents to Complete Large Software Tasks | Factory.ai
factory.ai · 5,404 words · saved by 1 readers
Models have become very good at problems with compact, stable criteria for success. Much of the past year's progress in...
Go back By Factory Research, Theo Luan - August 27, 2026 - 10 minute read Research Share Held to a standard of completion they wrote themselves, agents rebuilt complex programs to near-parity. Models have become very good at problems with compact, stable criteria for success. Much of the past year's progress in mathematics and constrained optimization falls into this category - machine-checked proofs of long-open Erdős problems, gold-medal performance at the IMO, new bounds on decades-old combinatorial problems. The search space can be enormous, but the result can ultimately be judged…
saved by
related reading
- Demystifying evals for AI agents \ Anthropicanthropic.com
- Composer2.pdfcursor.com
- Effective harnesses for long-running agents \ Anthropicanthropic.com
- Towards self-driving codebases · Cursorcursor.com
- Shipping at Inference-Speed | Peter Steinbergersteipete.me
- FrontierSWEfrontierswe.com
- After Automation | Everyevery.to
- Agent swarms and the new model economics · Cursorcursor.com
- Measuring AI Ability to Complete Long Tasks - METRmetr.org
- Scaling long-running autonomous coding · Cursorcursor.com
- Notes on the Software Factorybenedict.dev
- Senior SWE-Benchsenior-swe-bench.snorkel.ai