Expenditure Horizon: Measuring Optimization Ability, with an Application to NanoGPT - METR
metr.org · 9,144 words · saved by 1 readers
We propose a measure of an AI agent’s optimization ability with an "expenditure horizon." We give an empirical illustration from the NanoGPT speedrun.
We propose a measure of an AI agent’s optimization ability with an “expenditure horizon.” We give an empirical illustration from the NanoGPT speedrun. One difficulty in measuring AI’s ability to accelerate AI R&D is accounting for token cost, experiment compute cost and human labor cost. If we can estimate performance as a function of cost for both humans and agents, we can measure the “expenditure horizon” as the point at which those curves cross: the budget at which humans become more cost-effective than AIs. We illustrate our method with data from the NanoGPT speedrun. We first estimate…
saved by
related reading
- Evidence on AI R&D Progress from NanoGPT - METRmetr.org
- Estimating the Productivity of an Autonomous AI Software Engineer | Cognitioncognition.ai
- AI Timelines — LessWronglesswrong.com
- My picture of the present in AI — LessWronglesswrong.com
- An Apple-Picking Model of AI R&D | Tom Cunningham – Tom Cunninghamtecunningham.github.io
- Economics and Transformative AI | Tom Cunningham – Tom Cunninghamtecunningham.github.io
- AI 2027ai-2027.com
- Metrics of Agent Ability - METRmetr.org
- AI progress is about to speed up | Epoch AIepoch.ai
- I. From GPT-4 to AGI: Counting the OOMs - SITUATIONAL AWARENESSsituational-awareness.ai
- AGI is still 30 years away — Ege Erdil & Tamay Besirogludwarkesh.com
- Measuring AI Ability to Complete Long Tasks - METRmetr.org