Humans Still Beat AI in the Long Horizon: Revisiting Test-Time Scaling in the Agent Era | Qiuyang Mang
Agents can spend test-time compute by trying, observing, and revising. We derive an Elo reference for repeated sampling, then show that in a 2022 two-week coding marathon, current agents plateau within 24 hours while top humans keep improving.
TL;DR. Agents can spend test-time compute by trying, observing, and revising, so we ask whether their gains come from a better internal strategy or from something close to repeated sampling. We derive a simple Elo reference line: repeated sampling is linear in log test-time compute. In a 2022 two-week coding marathon, current agents plateau within 24 hours, while top humans keep improving over the official two weeks. The takeaway is that humans still do much better long-horizon test-time adaptation, and agent strategies have a lot of room to improve. Agents Bring Intrinsic Test-Time Strategies
Explore this link on the map →saved by
related reading
- An Apple-Picking Model of AI R&D | Tom Cunningham – Tom Cunninghamtecunningham.github.io
- Noam Brown on X: "Implications of Large-Scale Test-Time Compute" / Xx.com
- Demystifying evals for AI agents \ Anthropicanthropic.com
- AI 2027ai-2027.com
- Composer2.pdfcursor.com
- Agent Evaluation: A Detailed Guidecameronrwolfe.substack.com
- Measuring AI Ability to Complete Long Tasks - METRmetr.org
- The Era of Experience Paper.pdfstorage.googleapis.com
- Effective harnesses for long-running agents \ Anthropicanthropic.com
- We spent 2 hours working in the future - METRmetr.org
- Frontier Coding Agents Can Now Implement an AlphaZero Self-Play Machine Learning Pipeline For Connect Four That Performs Comparably to an External Solver — LessWronglesswrong.com
- AI 2027ai-2027.com