✳flâneur — a map of the web's best reading
[Interim research report] Taking features out of superposition with sparse autoencoders — LessWrong
lesswrong.com · 11,251 words · saved by 1 readers
We're thankful for helpful comments from Trenton Bricken, Eric Winsor, Noa Nabeshima, and Sid Black. …
x [Interim research report] Taking features out of superposition with sparse autoencoders — LessWrong Sparse Autoencoders (SAEs) Interpretability (ML & AI) Superposition Conjecture (org) Bucket Errors AI Frontpage 156 [Interim research report] Taking features out of superposition with sparse autoencoders by Lee Sharkey , Dan Braun , beren 13th Dec 2022 AI Alignment Forum 27 min read 23 156 Ω 69 We're thankful for helpful comments from Trenton Bricken, Eric Winsor, Noa Nabeshima, and Sid Black. This post is part of the work done at Conjecture . TL;DR : Recent results from Anthropic suggest that
Explore this link on the map →related reading
- Towards Monosemanticity: Decomposing Language Models With Dictionary Learningtransformer-circuits.pub
- Sparse Autoencoders Find Highly Interpretable Features in Language Modelsarxiv.org
- Toy Models of Superpositiontransformer-circuits.pub
- An Intuitive Explanation of Sparse Autoencoders for LLM Interpretability | Adam Karvonenadamkarvonen.github.io
- Open Source Replication & Commentary on Anthropic's Dictionary Learning Paper — LessWronglesswrong.com
- A gentle introduction to sparse autoencodersnickjiang.substack.com
- Do sparse autoencoders find "true features"? — LessWronglesswrong.com
- Circuits Updates - February 2024transformer-circuits.pub
- [2406.04093] Scaling and evaluating sparse autoencodersarxiv.org
- A gentle introduction to sparse autoencoders — LessWronglesswrong.com
- Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnettransformer-circuits.pub
- [2410.12101] The Persian Rug: solving toy models of superposition using large-scale symmetriesarxiv.org