✳flâneur — a map of the web's best reading
Noam Brown on X: "Implications of Large-Scale Test-Time Compute" / X
x.com · 630 words · saved by 7 readers
https://t.co/oWqzT12RtZ
@polynoamial: Implications of Large-Scale Test-Time Compute tl;dr: As LLMs become more capable, benchmark performance is increasingly a function of test-time compute. In fact, we likely don't know what the capability ceiling is for modern LLMs because it's too expensive to measure. We should change LLM evaluations to account for that by measuring performance vs tokens, cost, or time. The day GPT-5.5 was released, the initial reaction was skepticism. The benchmark numbers were better, but not by much: However, within hours, once people had time to play around with the model, it became clear th
Explore this link on the map →saved by
related reading
- Your Evals Will Break and You Won't See It Coming - Lun Wangwanglun1996.github.io
- My picture of the present in AI — LessWronglesswrong.com
- AI in 2025: gestalt — LessWronglesswrong.com
- Optimizing LLM Test-Time Compute Involves Solving a Meta-RL Problem – Machine Learning Blog | ML@CMU | Carnegie Mellon Universityblog.ml.cmu.edu
- Humans Still Beat AI in the Long Horizon: Revisiting Test-Time Scaling in the Agent Era | Qiuyang Mangjoyemang33.github.io
- gpt-4.pdfcdn.openai.com
- PostTrainBenchposttrainbench.com
- LLM Inference Performance Engineering: Best Practices | Databricks Blogdatabricks.com
- Are AI benchmarks doomed? - by Anson Ho and Greg Burnhamepochai.substack.com
- The bitter lesson of LLM evalsparsed.com
- I. From GPT-4 to AGI: Counting the OOMs - SITUATIONAL AWARENESSsituational-awareness.ai
- AI Model & API Providers Analysis | Artificial Analysisartificialanalysis.ai