k-Sparse Autoencoders
arxiv.org · 5,767 words · saved by 1 readers
N/A
k -Sparse Autoencoders Alireza Makhzani makhzani@psi.utoronto.ca Brendan Frey frey@psi.utoronto.ca University of Toronto, 10 King’s College Rd. Toronto, Ontario M5S 3G4, Canada arXiv:1312.5663v2 [cs.LG] 22 Mar 2014 Abstract and updating the dictionary.…
saved by
related reading
- An Intuitive Explanation of Sparse Autoencoders for LLM Interpretability | Adam Karvonenadamkarvonen.github.io
- [2608.27540] Towards a mathematical theory of superpositionarxiv.org
- sparseAutoencoder.pdfweb.stanford.edu
- Towards Monosemanticity: Decomposing Language Models With Dictionary Learningtransformer-circuits.pub
- Towards Monosemanticity: Decomposing Language Models With Dictionary Learningtransformer-circuits.pub
- sparse-autoencoders.pdfcdn.openai.com
- OpenAI SAE Training Paperarxiv.org
- [Interim research report] Taking features out of superposition with sparse autoencoders — LessWronglesswrong.com
- [2406.04093] Scaling and evaluating sparse autoencodersarxiv.org
- Sparse Autoencoders Find Highly Interpretable Features in Language Modelsarxiv.org
- Improving Dictionary Learning with Gated Sparse Autoencodersarxiv.org
- Autoencoder - Wikipediaen.wikipedia.org