Natural Language Autoencoders \ Anthropic
anthropic.com · 1,740 words · saved by 9 readers
Turning Claude's thoughts into text
Interpretability Natural Language Autoencoders: Turning Claude’s thoughts into text May 7, 2026 Read the paper When you talk to an AI model like Claude, you talk to it in words. Internally, Claude processes those words as long lists of numbers, before again producing words as its output. These numbers in the middle are called activations— and like neural activity in the human brain, they encode Claude’s thoughts. Also like neural activity, activations are difficult to understand. We can’t easily decode them to read Claude’s thoughts. Over the past few years, we’ve developed a range of tools (l
saved by
related reading
- Natural Language Autoencoders Produce Unsupervised Explanations of LLM Activationstransformer-circuits.pub
- Tracing the thoughts of a large language model \ Anthropicanthropic.com
- Natural Language Autoencoders Produce Unsupervised Explanations of LLM Activations — LessWronglesswrong.com
- On the Biology of a Large Language Modeltransformer-circuits.pub
- Transformer Circuits Threadtransformer-circuits.pub
- Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnettransformer-circuits.pub
- How AI Is Learning to Think in Secretnickandresen.substack.com
- Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnettransformer-circuits.pub
- A global workspace in language models \ Anthropicanthropic.com
- Neuronpedianeuronpedia.org
- GitHub - kitft/natural_language_autoencoders · GitHubgithub.com
- Emergent introspective awareness in large language models \ Anthropicanthropic.com