Is Capability a Liability? More Capable Language Models Make Worse Forecasts When It Matters Most
We document inverse scaling in LLMs on forecasting problems whose underlying time series exhibit superlinear growth and tail risk of regime change, a structure common in finance and epidemiology. On these tasks, more capable models produce worse distributional forecasts. The pattern appears on ForecastBench-Sim (FBSim), a contamination-free, simulated-world benchmark we release, in forecasting synthetic SIR epidemics with a matched linear control, and replicates in real-world datasets on COVID-19, measles, housing markets, and hyperinflation. A per-quantile decomposition shows the failure concentrates at the upper tail, which more capable models shift upward to track aggressive extrapolations of growth, while the lower tail stays put. A within-family study of Llama-3.1 shows that both model scale and post-training independently contribute to this effect. Domain knowledge does not reliably rescue calibration. This inverse scaling does not appear on single-threshold metrics common in LLM
Is Capability a Liability? More Capable Language Models Make Worse Forecasts When It Matters Most Nick Merrill Forecasting Research Institute UC Berkeley nick@forecastingresearch.org &Jaeho Lee Forecasting Research Institute jaeho@forecastingresearch.org &Ezra Karger Forecasting Research Institute ezra@forecastingresearch.org Abstract We document inverse scaling in LLMs on forecasting problems whose underlying time series exhibit superlinear growth and tail risk of regime change, a structure common in finance and epidemiology. On these tasks, more capable models produce worse distributional fo
Explore this link on the map →saved by
related reading
- Pitfalls in Evaluating Language Model Forecastersarxiv.org
- Your Evals Will Break and You Won't See It Coming - Lun Wangwanglun1996.github.io
- gpt-4.pdfcdn.openai.com
- AI in 2025: gestalt — LessWronglesswrong.com
- irmckenzie.co.uk/round2irmckenzie.co.uk
- [2405.10938] Observational Scaling Laws and the Predictability of Language Model Performancearxiv.org
- Training LLMs to Predict World Events (Guest Post with Mantic) - Thinking Machines Labthinkingmachines.ai
- How fast is AI improving? - AI Digesttheaidigest.org
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- larger language models may disappoint you [or, an eternally unfinished draft] — LessWronglesswrong.com
- [2603.02202] Frontier Models Can Take Actions at Low Probabilitiesarxiv.org
- The bitter lesson of LLM evalsparsed.com