Self-explaining SAE features — AI Alignment Forum
alignmentforum.org · 3,386 words · saved by 1 readers
TL;DR * We apply the method of SelfIE/Patchscopes to explain SAE features – we give the model a prompt like “What does X mean?”, replace the residua…
x Self-explaining SAE features — AI Alignment Forum Interpretability (ML & AI) MATS Program Sparse Autoencoders (SAEs) AI Frontpage 25 Self-explaining SAE features by Dmitrii Kharlapenko , neverix , Neel Nanda , Arthur Conmy 5th Aug 2024 12 min read 13 25 TL;DR We apply the method of SelfIE / Patchscopes to explain SAE features – we give the model a prompt like “What does X mean?”, replace the residual stream on X with the decoder direction times some scale, and have it generate an explanation. We call this self-explanation. The natural alternative is auto-interp, using a larger LLM to spot pa
related reading
- An Intuitive Explanation of Sparse Autoencoders for LLM Interpretability | Adam Karvonenadamkarvonen.github.io
- Open Source Automated Interpretability for Sparse Autoencoder Features | EleutherAI Blogblog.eleuther.ai
- A gentle introduction to sparse autoencoders — LessWronglesswrong.com
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnettransformer-circuits.pub
- Towards Monosemanticity: Decomposing Language Models With Dictionary Learningtransformer-circuits.pub
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- Prism: mapping interpretable concepts and features in a latent space of language | thesephist.comthesephist.com
- GitHub - PaulPauls/llama3_interpretability_sae: A complete end-to-end pipeline for LLM interpretability with sparse autoencoders (SAEs) using Llama 3.2, written in pure PyTorch and fully reproducible.github.com
- Neuronpedianeuronpedia.org
- A Survey on Sparse Autoencoders: Interpreting the Internal Mechanisms of Large Language Modelsarxiv.org
- Towards Monosemanticity: Decomposing Language Models With Dictionary Learningtransformer-circuits.pub