flâneur — a map of the web's best reading

Test your best methods on our hard CoT interp tasks — LessWrong

lesswrong.com · 13,754 words · saved by 1 readers

Daria and Riya are co-first authors. This work was done during Neel Nanda’s MATS 9.0. Claude helped write code and suggest edits for this post. Most of our tasks fall in 3 categories: predicting future actions, detecting the effect of an intervention, and identifying distributional properties of a rollout.[1] Baseline method results ordered by median OOD performance (g-mean²) across the 7 main tasks. Non-LLM methods have a slight lead, with TF-IDF and attention probes scoring significantly above chance. Few-shot confidence monitors, which receive both tuning and examples, perform best out of the LLM-based methods. Average g-mean² across 7 tasks Predicting what a model will do next (1) Will the model stop thinking soon? [2] (2) Will Gemma delete itself? (3) What will the model say if asked this question? Detecting whether a model was influenced by an intervention (4) Does the model really agree with the user, or is it being sycophantic? [3] (5) Did the model independently get to the ans

x Test your best methods on our hard CoT interp tasks — LessWrong Interpretability (ML & AI) MATS Program AI Frontpage 59 Test your best methods on our hard CoT interp tasks by daria , Riya Tyagi , Josh Engels , Neel Nanda 26th Mar 2026 AI Alignment Forum 23 min read 2 59 Ω 21 Authors: Daria Ivanova, Riya Tyagi, Josh Engels, Neel Nanda Daria and Riya are co-first authors. This work was done during Neel Nanda’s MATS 9.0. Claude helped write code and suggest edits for this post. Most of our tasks fall in 3 categories: predicting future actions, detecting the effect of an intervention, and identi

Explore this link on the map →

related reading