Natural Language Autoencoders Produce Unsupervised Explanations of LLM Activations — LessWrong
Abstract > We introduce Natural Language Autoencoders (NLAs), an unsupervised method for generating natural language explanations of LLM activations.…
x Natural Language Autoencoders Produce Unsupervised Explanations of LLM Activations — LessWrong AI Frontpage 2026 Top Fifty: 37 % 213 Natural Language Autoencoders Produce Unsupervised Explanations of LLM Activations by Subhash Kantamneni , kitft , Euan Ong , Sam Marks 7th May 2026 AI Alignment Forum 10 min read 35 213 Ω 77 Abstract We introduce Natural Language Autoencoders (NLAs), an unsupervised method for generating natural language explanations of LLM activations. An NLA consists of two LLM modules: an activation verbalizer (AV) that maps an activation to a text description and an activa
saved by
related reading
- Natural Language Autoencoders Produce Unsupervised Explanations of LLM Activationstransformer-circuits.pub
- Natural Language Autoencoders \ Anthropicanthropic.com
- GitHub - kitft/natural_language_autoencoders · GitHubgithub.com
- GitHub - asherps/EasyNLA: Minimal codebase for efficiently training Natural Language Autoencoders (NLAs). Built on Celeste's nanoNLA: https://github.com/ceselder/nanoNLAgithub.com
- Tracing the thoughts of a large language model \ Anthropicanthropic.com
- On the Biology of a Large Language Modeltransformer-circuits.pub
- Neuronpedianeuronpedia.org
- [2511.08579] Training Language Models to Explain Their Own Computationsarxiv.org
- Transformer Circuits Threadtransformer-circuits.pub
- Training Language Models to Explain Their Own Computationsarxiv.org
- Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnettransformer-circuits.pub
- Can activation verbalizers surface an internal chain of thought? — LessWronglesswrong.com