kitft/natural_language_autoencoders ·
github.com · 1,251 words · saved by 1 readers
No description, website, or topics provided.
Natural Language Autoencoders (NLA) Open-source library accompanying the Anthropic Transformer Circuits post Natural Language Autoencoders Produce Unsupervised Explanations of LLM Activations . 📄 Blog post · ▶ Video walkthrough · 🔬 Try the released NLAs on Neuronpedia A Natural Language Autoencoder is a pair of fine-tuned LMs that map residual-stream activation vectors to natural language and back: direction mechanism AV (activation verbalizer) vector → text inject the vector as a single token embedding into a fixed prompt, autoregress a description AR (activation reconstructor) text → vecto
related reading
- Natural Language Autoencoders Produce Unsupervised Explanations of LLM Activationstransformer-circuits.pub
- Natural Language Autoencoders Produce Unsupervised Explanations of LLM Activations — LessWronglesswrong.com
- GitHub - asherps/EasyNLA: Minimal codebase for efficiently training Natural Language Autoencoders (NLAs). Built on Celeste's nanoNLA: https://github.com/ceselder/nanoNLAgithub.com
- Natural Language Autoencoders \ Anthropicanthropic.com
- Towards Monosemanticity: Decomposing Language Models With Dictionary Learningtransformer-circuits.pub
- An Intuitive Explanation of Sparse Autoencoders for LLM Interpretability | Adam Karvonenadamkarvonen.github.io
- Transformer Circuits Threadtransformer-circuits.pub
- Neuronpedianeuronpedia.org
- Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnettransformer-circuits.pub
- GitHub - PaulPauls/llama3_interpretability_sae: A complete end-to-end pipeline for LLM interpretability with sparse autoencoders (SAEs) using Llama 3.2, written in pure PyTorch and fully reproducible.github.com
- I Trained a Language Model. Then I Built a Brain Scanner and Looked Inside It. | by Caleb DeLeeuw | Mediummedium.com
- Towards Monosemanticity: Decomposing Language Models With Dictionary Learningtransformer-circuits.pub