✳flâneur — a map of the web's best reading
Can activation verbalizers surface an internal chain of thought? — LessWrong
lesswrong.com · 11,464 words · saved by 4 readers
We introduce an evaluation for activation verbalizers: can they surface a target model's reasoning as it solves a math problem in a single forward pa…
x Can activation verbalizers surface an internal chain of thought? — LessWrong Interpretability (ML & AI) AI Frontpage 2026 Top Fifty: 14 % 122 Can activation verbalizers surface an internal chain of thought? by oakhu , ryan_greenblatt 7th Jun 2026 AI Alignment Forum 19 min read 0 122 Ω 53 We introduce an evaluation for activation verbalizers: can they surface a target model's reasoning as it solves a math problem in a single forward pass? For open-weight NLAs, the answer seems to be: "possibly, but definitely not reliably". Lots of important capabilities currently require AI models to reason
Explore this link on the map →saved by
related reading
- Prediction, Explanation, or Over-interpretation?elena-baixy.github.io
- [2510.24941] Can Aha Moments Be Fake? Towards Quantifying Decorative and True Thinking in Chain-of-Thoughtarxiv.org
- Natural Language Autoencoders Produce Unsupervised Explanations of LLM Activationstransformer-circuits.pub
- Verbalizable Representations Form a Global Workspace in Language Modelstransformer-circuits.pub
- Tracing the thoughts of a large language model \ Anthropicanthropic.com
- [2201.11903] Chain of Thought Prompting Elicits Reasoning in Large Language Modelsarxiv.org
- Natural Language Autoencoders \ Anthropicanthropic.com
- Externalized reasoning oversight: a research direction for language model alignment — AI Alignment Forumalignmentforum.org
- The Unintelligibility is Ours: Notes on Chain of Thought1a3orn.com
- How do we know if activation verbalizers are telling us anything about activations? — Millicent Limillicentli.github.io
- the-illusion-of-thinking.pdfml-site.cdn-apple.com
- Natural Language Autoencoders Produce Unsupervised Explanations of LLM Activations — LessWronglesswrong.com