✳flâneur — a map of the web's best reading
The Geometry of Concepts: Sparse Autoencoder Feature Structure
arxiv.org · saved by 1 readers
N/A
Explore this link on the map →related reading
- Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnettransformer-circuits.pub
- Towards Monosemanticity: Decomposing Language Models With Dictionary Learningtransformer-circuits.pub
- Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnettransformer-circuits.pub
- sparse-autoencoders.pdfcdn.openai.com
- An Intuitive Explanation of Sparse Autoencoders for LLM Interpretability | Adam Karvonenadamkarvonen.github.io
- Matryoshka Sparse Autoencoders — LessWronglesswrong.com
- GitHub - PaulPauls/llama3_interpretability_sae: A complete end-to-end pipeline for LLM interpretability with sparse autoencoders (SAEs) using Llama 3.2, written in pure PyTorch and fully reproducible. · GitHubgithub.com
- Neuronpedianeuronpedia.org
- Sparse Autoencoders Find Highly Interpretable Features in Language Modelsarxiv.org
- Symmetry and Geometry in Neural Representationsproceedings.mlr.press
- pdfopenreview.net
- Do sparse autoencoders find "true features"? — LessWronglesswrong.com