flâneur — a map of the web's best reading

Measuring no CoT math time horizon (single forward pass)

blog.redwoodresearch.org · 1,444 words · saved by 1 readers

Opus 4.5 has around a 3.5 minute 50%-reliablity time horizon

Measuring no CoT math time horizon (single forward pass) Opus 4.5 has around a 3.5 minute 50%-reliablity time horizon Ryan Greenblatt Dec 26, 2025 11 1 Share A key risk factor for scheming (and misalignment more generally) is opaque reasoning ability . One proxy for this is how good AIs are at solving math problems immediately without any chain-of-thought (CoT) (as in, in a single forward pass). I’ve measured this on a dataset of easy math problems and used this to estimate 50% reliability no-CoT time horizon using the same methodology introduced in Measuring AI Ability to Complete Long Tasks

Explore this link on the map →

saved by

related reading