Natural Language Autoencoders Produce Unsupervised Explanations of LLM Activations
We introduce Natural Language Autoencoders (NLAs), an unsupervised method for generating natural language explanations of LLM activations. An NLA consists of two LLM modules: an activation verbalizer (AV) that maps an activation to a text description and an activation reconstructor (AR) that maps the description back to an activation. We jointly train the AV and AR with reinforcement learning to reconstruct residual stream activations. Although we optimize for activation reconstruction, the resulting NLA explanations read as plausible interpretations of model internals that, according to our quantitative evaluations, grow more informative over training.
Natural Language Autoencoders Produce Unsupervised Explanations of LLM Activations Transformer Circuits Thread Natural Language Autoencoders Produce Unsupervised Explanations of LLM Activations Natural Language Autoencoders Produce Unsupervised Explanations of LLM Activations Authors Kit Fraser-Taliente*, Subhash Kantamneni*‡, Euan Ong*, Dan Mossing, Christina Lu, Paul C. Bogdan Emmanuel Ameisen, James Chen, Dzmitry Kishylau, Adam Pearce, Julius Tarng, Alex Wu, Jeff Wu, Yang Zhang, Daniel M. Ziegler Evan Hubinger, Joshua Batson, Jack Lindsey, Samuel Zimmerman, Samuel Marks Affiliations Anthrop
Explore this link on the map →saved by
related reading
- Natural Language Autoencoders Produce Unsupervised Explanations of LLM Activations — LessWronglesswrong.com
- Natural Language Autoencoders \ Anthropicanthropic.com
- Neuronpedianeuronpedia.org
- GitHub - kitft/natural_language_autoencoders · GitHubgithub.com
- Tracing the thoughts of a large language model \ Anthropicanthropic.com
- On the Biology of a Large Language Modeltransformer-circuits.pub
- Can activation verbalizers surface an internal chain of thought? — LessWronglesswrong.com
- Transformer Circuits Threadtransformer-circuits.pub
- Dreaming Vectors: Gradient-descented steering vectors from Activation Oracles and using them to Red-Team AOs — LessWronglesswrong.com
- LLM Visualizationbbycroft.net
- Prediction, Explanation, or Over-interpretation?elena-baixy.github.io
- An Intuitive Explanation of Sparse Autoencoders for LLM Interpretability | Adam Karvonenadamkarvonen.github.io