✳flâneur — a map of the web's best reading
Measuring no CoT math time horizon (single forward pass)
blog.redwoodresearch.org · 1,444 words · saved by 1 readers
Opus 4.5 has around a 3.5 minute 50%-reliablity time horizon
Measuring no CoT math time horizon (single forward pass) Opus 4.5 has around a 3.5 minute 50%-reliablity time horizon Ryan Greenblatt Dec 26, 2025 11 1 Share A key risk factor for scheming (and misalignment more generally) is opaque reasoning ability . One proxy for this is how good AIs are at solving math problems immediately without any chain-of-thought (CoT) (as in, in a single forward pass). I’ve measured this on a dataset of easy math problems and used this to estimate 50% reliability no-CoT time horizon using the same methodology introduced in Measuring AI Ability to Complete Long Tasks
Explore this link on the map →saved by
related reading
- My picture of the present in AI — LessWronglesswrong.com
- the-illusion-of-thinking.pdfml-site.cdn-apple.com
- Mathematics in the Library of Babel - Daniel Littdaniellitt.com
- I. From GPT-4 to AGI: Counting the OOMs - SITUATIONAL AWARENESSsituational-awareness.ai
- Measuring AI Ability to Complete Long Tasks - METRmetr.org
- AI progress is about to speed up | Epoch AIepoch.ai
- How Does Time Horizon Vary Across Domains? - METRmetr.org
- Recent LLMs can do 2-hop and 3-hop latent (no-CoT) reasoning on natural facts — LessWronglesswrong.com
- o1 and Reasoning | AndoLogsblog.ando.ai
- Chain-of-Thought Promptinglearnprompting.org
- AI #100: Meet the New Boss | Don't Worry About the Vasethezvi.wordpress.com
- A recent experience with ChatGPT 5.5 Pro | Gowers's Webloggowers.wordpress.com