flâneur — a map of the web's best reading

Can activation verbalizers surface an internal chain of thought? — LessWrong

lesswrong.com · 11,464 words · saved by 4 readers

We introduce an evaluation for activation verbalizers: can they surface a target model's reasoning as it solves a math problem in a single forward pa…

x Can activation verbalizers surface an internal chain of thought? — LessWrong Interpretability (ML & AI) AI Frontpage 2026 Top Fifty: 14 % 122 Can activation verbalizers surface an internal chain of thought? by oakhu , ryan_greenblatt 7th Jun 2026 AI Alignment Forum 19 min read 0 122 Ω 53 We introduce an evaluation for activation verbalizers: can they surface a target model's reasoning as it solves a math problem in a single forward pass? For open-weight NLAs, the answer seems to be: "possibly, but definitely not reliably". Lots of important capabilities currently require AI models to reason

Explore this link on the map →

saved by

related reading