Noam Brown on X: "Implications of Large-Scale Test-Time Compute" / X
x.com · 630 words · saved by 7 readers
https://t.co/oWqzT12RtZ
@polynoamial: Implications of Large-Scale Test-Time Compute tl;dr: As LLMs become more capable, benchmark performance is increasingly a function of test-time compute. In fact, we likely don't know what the capability ceiling is for modern LLMs because it's too expensive to measure. We should change LLM evaluations to account for that by measuring performance vs tokens, cost, or time. The day GPT-5.5 was released, the initial reaction was skepticism. The benchmark numbers were better, but not by much: However, within hours, once people had time to play around with the model, it became clear th
saved by
related reading
- Your Evals Will Break and You Won't See It Coming - Lun Wangwanglun1996.github.io
- My picture of the present in AI — LessWronglesswrong.com
- More compute, more capability: Why AI agent evaluations need to account for test-time compute | AISI Workaisi.gov.uk
- AI in 2025: gestalt — LessWronglesswrong.com
- On measuring AInikilravi.substack.com
- Humans Still Beat AI in the Long Horizon: Revisiting Test-Time Scaling in the Agent Era | Qiuyang Mangjoyemang33.github.io
- Optimizing LLM Test-Time Compute Involves Solving a Meta-RL Problem – Machine Learning Blog | ML@CMU | Carnegie Mellon Universityblog.ml.cmu.edu
- gpt-4.pdfcdn.openai.com
- Are AI benchmarks doomed? - by Anson Ho and Greg Burnhamepochai.substack.com
- The bitter lesson of LLM evalsparsed.com
- I. From GPT-4 to AGI: Counting the OOMs - SITUATIONAL AWARENESSsituational-awareness.ai
- Late Takes on OpenAI o1alexirpan.com