flâneur — a map of the web's best reading

Noam Brown on X: "Implications of Large-Scale Test-Time Compute" / X

x.com · 630 words · saved by 7 readers

https://t.co/oWqzT12RtZ

@polynoamial: Implications of Large-Scale Test-Time Compute tl;dr: As LLMs become more capable, benchmark performance is increasingly a function of test-time compute. In fact, we likely don't know what the capability ceiling is for modern LLMs because it's too expensive to measure. We should change LLM evaluations to account for that by measuring performance vs tokens, cost, or time. The day GPT-5.5 was released, the initial reaction was skepticism. The benchmark numbers were better, but not by much: However, within hours, once people had time to play around with the model, it became clear th

Explore this link on the map →

saved by

related reading