flâneur

LLMs Are Closing the Gap on Human Superforecasters

substack.com · 539 words · saved by 1 readers

We opened our AI forecasting benchmark to external submissions. Here’s what happened.

(Post written by Houtan Bastani, Simas Kučinskas, and Matt Reynolds)In October, we opened our AI forecasting benchmark, ForecastBench, to external submissions. The challenge: beat superforecasters using any tools available, from scaffolding to fine-tuning. Several teams responded, including xAI, Cassi, Lightning Rod, and Mantic. We thank all of them for participating on this challenging benchmark. The result? External submissions now hold #2 and #3 on our leaderboard, outperforming all our baseline LLM configurations. Superforecasters still hold #1. Here are the current standings (lower…

saved by

related reading