flâneur — a map of the web's best reading

Thought Anchors: Which LLM Reasoning Steps Matter? — LessWrong

lesswrong.com · 2,506 words · saved by 1 readers

This post is adapted from our recent arXiv paper. Paul Bogdan and Uzay Macar are co-first authors on this work. Reasoning LLMs have recently achieved state-of-the-art performance across many domains. However, their long-form CoT reasoning creates significant interpretability challenges as each generated token depends on all previous ones, making the underlying computation difficult to decompose and understand. We present a vision for interpreting CoT by decomposing it into sentences. A sentence-level analysis is more manageable than looking at the token level, as sentences are usually 10-30x less numerous. Additionally, one sentence often puts forth a single proposition, leading us to suspect this is a meaningful logical unit. While future work may yield more sophisticated strategies for dividing CoT into atomic steps, we believe sentences can serve as a reasonable starting point for investigating multi-token computations. We find that not all sentences in these reasoning traces are cr

x Thought Anchors: Which LLM Reasoning Steps Matter? — LessWrong Chain-of-Thought Alignment Language Models (LLMs) MATS Program AI Frontpage 36 Thought Anchors: Which LLM Reasoning Steps Matter? by Uzay Macar , Paul B , Neel Nanda , Arthur Conmy 2nd Jul 2025 AI Alignment Forum Linkpost for www.thought-anchors.com 7 min read 6 36 Ω 16 This post is adapted from our recent arXiv paper . Paul Bogdan and Uzay Macar are co-first authors on this work. TL;DR Interpretability of chains-of-thought (CoTs) produced by LLMs is challenging: Standard mechanistic interpretability studies a single token's gene

Explore this link on the map →

related reading