✳flâneur — a map of the web's best reading
Self-explaining SAE features — AI Alignment Forum
alignmentforum.org · 3,386 words · saved by 1 readers
TL;DR * We apply the method of SelfIE/Patchscopes to explain SAE features – we give the model a prompt like “What does X mean?”, replace the residua…
x Self-explaining SAE features — AI Alignment Forum Interpretability (ML & AI) MATS Program Sparse Autoencoders (SAEs) AI Frontpage 25 Self-explaining SAE features by Dmitrii Kharlapenko , neverix , Neel Nanda , Arthur Conmy 5th Aug 2024 12 min read 13 25 TL;DR We apply the method of SelfIE / Patchscopes to explain SAE features – we give the model a prompt like “What does X mean?”, replace the residual stream on X with the decoder direction times some scale, and have it generate an explanation. We call this self-explanation. The natural alternative is auto-interp, using a larger LLM to spot pa
Explore this link on the map →related reading
- An Intuitive Explanation of Sparse Autoencoders for LLM Interpretability | Adam Karvonenadamkarvonen.github.io
- Open Source Automated Interpretability for Sparse Autoencoder Features | EleutherAI Blogblog.eleuther.ai
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnettransformer-circuits.pub
- Towards Monosemanticity: Decomposing Language Models With Dictionary Learningtransformer-circuits.pub
- A gentle introduction to sparse autoencoders — LessWronglesswrong.com
- Neuronpedianeuronpedia.org
- Prism: mapping interpretable concepts and features in a latent space of language | thesephist.comthesephist.com
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- pdfopenreview.net
- An Intuitive Explanation of Sparse Autoencoders for Mechanistic Interpretability of LLMs — LessWronglesswrong.com
- Interpretability with Sparse Autoencoders (Colab exercises) — LessWronglesswrong.com