LLMs Are Closing the Gap on Human Superforecasters
substack.com · 539 words · saved by 1 readers
We opened our AI forecasting benchmark to external submissions. Here’s what happened.
(Post written by Houtan Bastani, Simas Kučinskas, and Matt Reynolds)In October, we opened our AI forecasting benchmark, ForecastBench, to external submissions. The challenge: beat superforecasters using any tools available, from scaffolding to fine-tuning. Several teams responded, including xAI, Cassi, Lightning Rod, and Mantic. We thank all of them for participating on this challenging benchmark. The result? External submissions now hold #2 and #3 on our leaderboard, outperforming all our baseline LLM configurations. Superforecasters still hold #1. Here are the current standings (lower…
saved by
related reading
- A Professional Superforecaster Walks Us Through His AI Progress Forecastssubstack.com
- The AI Superforecasters Are Here - by Scott Alexanderastralcodexten.com
- Exploreforecastbench.org
- Pitfalls in Evaluating Language Model Forecastersarxiv.org
- Training LLMs to Predict World Events (Guest Post with Mantic) - Thinking Machines Labthinkingmachines.ai
- AI in 2025: gestalt — LessWronglesswrong.com
- near-term-xpt-accuracy.pdfstatic1.squarespace.com
- Is Capability a Liability? More Capable Language Models Make Worse Forecasts When It Matters Mostarxiv.org
- Superhuman Automated Forecasting | CAISsafe.ai
- 2025: The year in LLMssimonwillison.net
- What We’ve Learned From A Year of Building with LLMs – Applied LLMsapplied-llms.org
- PostTrainBenchposttrainbench.com