flâneur — a map of the web's best reading

Is Capability a Liability? More Capable Language Models Make Worse Forecasts When It Matters Most

arxiv.org · 20,176 words · saved by 1 readers

We document inverse scaling in LLMs on forecasting problems whose underlying time series exhibit superlinear growth and tail risk of regime change, a structure common in finance and epidemiology. On these tasks, more capable models produce worse distributional forecasts. The pattern appears on ForecastBench-Sim (FBSim), a contamination-free, simulated-world benchmark we release, in forecasting synthetic SIR epidemics with a matched linear control, and replicates in real-world datasets on COVID-19, measles, housing markets, and hyperinflation. A per-quantile decomposition shows the failure concentrates at the upper tail, which more capable models shift upward to track aggressive extrapolations of growth, while the lower tail stays put. A within-family study of Llama-3.1 shows that both model scale and post-training independently contribute to this effect. Domain knowledge does not reliably rescue calibration. This inverse scaling does not appear on single-threshold metrics common in LLM

Is Capability a Liability? More Capable Language Models Make Worse Forecasts When It Matters Most Nick Merrill Forecasting Research Institute UC Berkeley nick@forecastingresearch.org &Jaeho Lee Forecasting Research Institute jaeho@forecastingresearch.org &Ezra Karger Forecasting Research Institute ezra@forecastingresearch.org Abstract We document inverse scaling in LLMs on forecasting problems whose underlying time series exhibit superlinear growth and tail risk of regime change, a structure common in finance and epidemiology. On these tasks, more capable models produce worse distributional fo

Explore this link on the map →

saved by

related reading