Natural Language Autoencoders Produce Unsupervised Explanations of LLM Activations — LessWrong
Abstract > We introduce Natural Language Autoencoders (NLAs), an unsupervised method for generating natural language explanations of LLM activations.…
x Natural Language Autoencoders Produce Unsupervised Explanations of LLM Activations — LessWrong AI Frontpage 2026 Top Fifty: 37 % 213 Natural Language Autoencoders Produce Unsupervised Explanations of LLM Activations by Subhash Kantamneni , kitft , Euan Ong , Sam Marks 7th May 2026 AI Alignment Forum 10 min read 35 213 Ω 77 Abstract We introduce Natural Language Autoencoders (NLAs), an unsupervised method for generating natural language explanations of LLM activations. An NLA consists of two LLM modules: an activation verbalizer (AV) that maps an activation to a text description and an activa
Explore this link on the map →saved by
related reading
- Natural Language Autoencoders Produce Unsupervised Explanations of LLM Activationstransformer-circuits.pub
- Natural Language Autoencoders \ Anthropicanthropic.com
- GitHub - kitft/natural_language_autoencoders · GitHubgithub.com
- Tracing the thoughts of a large language model \ Anthropicanthropic.com
- On the Biology of a Large Language Modeltransformer-circuits.pub
- Neuronpedianeuronpedia.org
- Transformer Circuits Threadtransformer-circuits.pub
- Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnettransformer-circuits.pub
- Can activation verbalizers surface an internal chain of thought? — LessWronglesswrong.com
- Prediction, Explanation, or Over-interpretation?elena-baixy.github.io
- GitHub - PaulPauls/llama3_interpretability_sae: A complete end-to-end pipeline for LLM interpretability with sparse autoencoders (SAEs) using Llama 3.2, written in pure PyTorch and fully reproducible. · GitHubgithub.com
- Activation Oracles: Training and Evaluating LLMs as General-Purpose Activation Explainers — LessWronglesswrong.com