✳flâneur — a map of the web's best reading
Matryoshka Sparse Autoencoders — LessWrong
lesswrong.com · 6,247 words · saved by 2 readers
View trees here Search through latents with a token-regex language View individual latents here See code here (github.com/noanabeshima/matryoshka-sae…
x Matryoshka Sparse Autoencoders — LessWrong Interpretability (ML & AI) Sparse Autoencoders (SAEs) AI Frontpage 100 Matryoshka Sparse Autoencoders by Noa Nabeshima 14th Dec 2024 AI Alignment Forum 13 min read 15 100 Ω 51 View trees here Search through latents with a token-regex language View individual latents here See code here (github.com/noanabeshima/matryoshka-saes) Alternate version of this document with appropriate-height interactives. Abstract Sparse autoencoders (SAEs) [1] [2] break down neural network internals into components called latents. Smaller SAE latents seem to correspond to
Explore this link on the map →saved by
related reading
- Towards Monosemanticity: Decomposing Language Models With Dictionary Learningtransformer-circuits.pub
- pdfopenreview.net
- An Intuitive Explanation of Sparse Autoencoders for LLM Interpretability | Adam Karvonenadamkarvonen.github.io
- [2409.14507] A is for Absorption: Studying Feature Splitting and Absorption in Sparse Autoencodersar5iv.labs.arxiv.org
- A gentle introduction to sparse autoencoders — LessWronglesswrong.com
- Do sparse autoencoders find "true features"? — LessWronglesswrong.com
- A gentle introduction to sparse autoencodersnickjiang.substack.com
- Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnettransformer-circuits.pub
- Identifying Functionally Important Features with End-to-End Sparse Dictionary Learning — LessWronglesswrong.com
- An Intuitive Explanation of Sparse Autoencoders for Mechanistic Interpretability of LLMs — LessWronglesswrong.com
- Interpretability with Sparse Autoencoders (Colab exercises) — LessWronglesswrong.com
- Sparse Autoencoders Find Highly Interpretable Features in Language Modelsarxiv.org