✳flâneur — a map of the web's best reading
kitft/natural_language_autoencoders ·
github.com · 1,251 words · saved by 1 readers
No description, website, or topics provided.
Natural Language Autoencoders (NLA) Open-source library accompanying the Anthropic Transformer Circuits post Natural Language Autoencoders Produce Unsupervised Explanations of LLM Activations . 📄 Blog post · ▶ Video walkthrough · 🔬 Try the released NLAs on Neuronpedia A Natural Language Autoencoder is a pair of fine-tuned LMs that map residual-stream activation vectors to natural language and back: direction mechanism AV (activation verbalizer) vector → text inject the vector as a single token embedding into a fixed prompt, autoregress a description AR (activation reconstructor) text → vecto
Explore this link on the map →related reading
- Natural Language Autoencoders Produce Unsupervised Explanations of LLM Activationstransformer-circuits.pub
- Natural Language Autoencoders Produce Unsupervised Explanations of LLM Activations — LessWronglesswrong.com
- Natural Language Autoencoders \ Anthropicanthropic.com
- Towards Monosemanticity: Decomposing Language Models With Dictionary Learningtransformer-circuits.pub
- Neuronpedianeuronpedia.org
- Transformer Circuits Threadtransformer-circuits.pub
- An Intuitive Explanation of Sparse Autoencoders for LLM Interpretability | Adam Karvonenadamkarvonen.github.io
- Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnettransformer-circuits.pub
- I Trained a Language Model. Then I Built a Brain Scanner and Looked Inside It. | by Caleb DeLeeuw | Mediummedium.com
- Transformer Explainer: LLM Transformer Model Visually Explainedpoloclub.github.io
- GenAI Handbookgenai-handbook.github.io
- Jacobian Lens – Qwen3.6-27B | Neuronpedianeuronpedia.org