OpenAI SAE Training Paper
arxiv.org · 6,903 words · saved by 1 readers
N/A
Scaling and evaluating sparse autoencoders Leo Gao∗ Tom Dupré la Tour† Henk Tillman† Gabriel Goh Rajan Troll Alec Radford Ilya Sutskever Jan Leike Jeffrey Wu† arXiv:2406.04093v1 [cs.LG] 6 Jun 2024 OpenAI…
related reading
- sparse-autoencoders.pdfcdn.openai.com
- [2406.04093] Scaling and evaluating sparse autoencodersarxiv.org
- An Intuitive Explanation of Sparse Autoencoders for LLM Interpretability | Adam Karvonenadamkarvonen.github.io
- A gentle introduction to sparse autoencoders — LessWronglesswrong.com
- Towards Monosemanticity: Decomposing Language Models With Dictionary Learningtransformer-circuits.pub
- Sparse Autoencoders Find Highly Interpretable Features in Language Modelsarxiv.org
- Matryoshka Sparse Autoencoders — LessWronglesswrong.com
- A Survey on Sparse Autoencoders: Interpreting the Internal Mechanisms of Large Language Modelsarxiv.org
- pdfopenreview.net
- Llama Scope: Extracting Features from Llama 3.1-8B with SAEsarxiv.org
- A gentle introduction to sparse autoencodersnickjiang.substack.com
- [Interim research report] Taking features out of superposition with sparse autoencoders — LessWronglesswrong.com