2503.14499
arxiv.org · 8,223 words · saved by 1 readers
N/A
Measuring AI Ability to Complete Long Software Tasks Thomas Kwa∗†, Ben West∗, Joel Becker, Amy Deng, Katharyn Garcia, Max Hasin, Sami Jawhar, Megan Kinniment, Nate Rush, Sydney Von Arx Ryan Bloom, Thomas Broadley, Haoxing Du, Brian Goodrich, Nikola Jurkovic, Luke Harold Miles‡, Seraphina Nix, Tao Lin, Chris…
related reading
- Measuring AI Ability to Complete Long Software Tasksarxiv.org
- Measuring AI Ability to Complete Long Tasks - METRmetr.org
- My picture of the present in AI — LessWronglesswrong.com
- I underestimated AI capabilities (again) - by Ajeya Cotraplanned-obsolescence.org
- Clarifying and predicting AGI — LessWronglesswrong.com
- A (Slightly) Mechanistic Theory for Exponentially Increasing AI Time Horizons? — LessWronglesswrong.com
- Where’s my ten minute AGI?epochai.substack.com
- AI in 2025: gestalt — LessWronglesswrong.com
- AI Timelines — LessWronglesswrong.com
- How Does Time Horizon Vary Across Domains? - METRmetr.org
- AI progress is about to speed up | Epoch AIepoch.ai
- Are AI benchmarks doomed? - by Anson Ho and Greg Burnhamepochai.substack.com