sparse-autoencoders.pdf
cdn.openai.com · 7,119 words · saved by 1 readers
N/A
Scaling and evaluating sparse autoencoders Leo Gao∗ Tom Dupré la Tour† Henk Tillman† Gabriel Goh Rajan Troll Alec Radford Ilya Sutskever Jan Leike Jeffrey Wu† OpenAI Abstract Sparse autoencoders provide a promising unsupervised approach for extracting in- terpretable features from a language model by…
related reading
- OpenAI SAE Training Paperarxiv.org
- [2406.04093] Scaling and evaluating sparse autoencodersarxiv.org
- An Intuitive Explanation of Sparse Autoencoders for LLM Interpretability | Adam Karvonenadamkarvonen.github.io
- Sparse Autoencoders Find Highly Interpretable Features in Language Modelsarxiv.org
- Towards Monosemanticity: Decomposing Language Models With Dictionary Learningtransformer-circuits.pub
- A gentle introduction to sparse autoencoders — LessWronglesswrong.com
- Matryoshka Sparse Autoencoders — LessWronglesswrong.com
- [Interim research report] Taking features out of superposition with sparse autoencoders — LessWronglesswrong.com
- A Survey on Sparse Autoencoders: Interpreting the Internal Mechanisms of Large Language Modelsarxiv.org
- A gentle introduction to sparse autoencodersnickjiang.substack.com
- pdfopenreview.net
- Towards Monosemanticity: Decomposing Language Models With Dictionary Learningtransformer-circuits.pub