Humans Still Beat AI in the Long Horizon: Revisiting Test-Time Scaling in the Agent Era | Qiuyang Mang
Agents can spend test-time compute by trying, observing, and revising. We derive an Elo reference for repeated sampling, then show that in a 2022 two-week coding marathon, current agents plateau within 24 hours while top humans keep improving.
TL;DR. Agents can spend test-time compute by trying, observing, and revising, so we ask whether their gains come from a better internal strategy or from something close to repeated sampling. We derive a simple Elo reference line: repeated sampling is linear in log test-time compute. In a 2022 two-week coding marathon, current agents plateau within 24 hours, while top humans keep improving over the official two weeks. The takeaway is that humans still do much better long-horizon test-time adaptation, and agent strategies have a lot of room to improve. Agents Bring Intrinsic Test-Time Strategies
saved by
related reading
- Machine Studying | Jacob Xiaochen Lijacobxli.com
- An Apple-Picking Model of AI R&D | Tom Cunningham – Tom Cunninghamtecunningham.github.io
- Demystifying evals for AI agents \ Anthropicanthropic.com
- Composer2.pdfcursor.com
- Measuring AI Ability to Complete Long Tasks - METRmetr.org
- More compute, more capability: Why AI agent evaluations need to account for test-time compute | AISI Workaisi.gov.uk
- Agent Evaluation: A Detailed Guidecameronrwolfe.substack.com
- Effective harnesses for long-running agents \ Anthropicanthropic.com
- We spent 2 hours working in the future - METRmetr.org
- Clarifying and predicting AGI — LessWronglesswrong.com
- EdgeBench | Scaling Laws of Environment Learningedge-bench.org
- After Automation | Everyevery.to