Test your best methods on our hard CoT interp tasks — LessWrong
Daria and Riya are co-first authors. This work was done during Neel Nanda’s MATS 9.0. Claude helped write code and suggest edits for this post. Most of our tasks fall in 3 categories: predicting future actions, detecting the effect of an intervention, and identifying distributional properties of a rollout.[1] Baseline method results ordered by median OOD performance (g-mean²) across the 7 main tasks. Non-LLM methods have a slight lead, with TF-IDF and attention probes scoring significantly above chance. Few-shot confidence monitors, which receive both tuning and examples, perform best out of the LLM-based methods. Average g-mean² across 7 tasks Predicting what a model will do next (1) Will the model stop thinking soon? [2] (2) Will Gemma delete itself? (3) What will the model say if asked this question? Detecting whether a model was influenced by an intervention (4) Does the model really agree with the user, or is it being sycophantic? [3] (5) Did the model independently get to the ans
x Test your best methods on our hard CoT interp tasks — LessWrong Interpretability (ML & AI) MATS Program AI Frontpage 59 Test your best methods on our hard CoT interp tasks by daria , Riya Tyagi , Josh Engels , Neel Nanda 26th Mar 2026 AI Alignment Forum 23 min read 2 59 Ω 21 Authors: Daria Ivanova, Riya Tyagi, Josh Engels, Neel Nanda Daria and Riya are co-first authors. This work was done during Neel Nanda’s MATS 9.0. Claude helped write code and suggest edits for this post. Most of our tasks fall in 3 categories: predicting future actions, detecting the effect of an intervention, and identi
Explore this link on the map →related reading
- Tracing the thoughts of a large language model \ Anthropicanthropic.com
- Thought Branches: Interpreting LLM Reasoning Requires Resamplingarxiv.org
- confessions_paper.pdfcdn.openai.com
- When Chain of Thought is Necessary, Language Models Struggle to Evade Monitorsarxiv.org
- the case for CoT unfaithfulness is overstated — LessWronglesswrong.com
- Chain-of-Thought Promptinglearnprompting.org
- [2510.24941] Can Aha Moments Be Fake? Towards Quantifying Decorative and True Thinking in Chain-of-Thoughtarxiv.org
- Can activation verbalizers surface an internal chain of thought? — LessWronglesswrong.com
- Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscationarxiv.org
- [2505.05410] Reasoning Models Don't Always Say What They Thinkarxiv.org
- Measuring Faithfulness in Chain-of-Thought Reasoning \ Anthropicanthropic.com
- Towards Faithful Chain-of-Thought: Large Language Models are Bridging Reasonersarxiv.org