✳flâneur — a map of the web's best reading
Addressing Feature Suppression in SAEs — LessWrong
lesswrong.com · 3,498 words · saved by 1 readers
Produced as part of the ML Alignment Theory Scholars Program - Winter 2023-24 Cohort as part of Lee Sharkey's stream. …
x Addressing Feature Suppression in SAEs — LessWrong Sparse Autoencoders (SAEs) Interpretability (ML & AI) AI Frontpage 88 Addressing Feature Suppression in SAEs by Benjamin Wright , Lee Sharkey 16th Feb 2024 AI Alignment Forum 12 min read 5 88 Ω 46 Produced as part of the ML Alignment Theory Scholars Program - Winter 2023-24 Cohort as part of Lee Sharkey's stream. TL;DR Sparse autoencoders are a method of resolving superposition by recovering linearly encoded “features” inside activations. Unfortunately, despite the great recent success of SAEs at extracting human interpretable features, they
Explore this link on the map →related reading
- An Intuitive Explanation of Sparse Autoencoders for LLM Interpretability | Adam Karvonenadamkarvonen.github.io
- pdfopenreview.net
- A gentle introduction to sparse autoencoders — LessWronglesswrong.com
- Do sparse autoencoders find "true features"? — LessWronglesswrong.com
- Identifying Functionally Important Features with End-to-End Sparse Dictionary Learning — LessWronglesswrong.com
- Towards Monosemanticity: Decomposing Language Models With Dictionary Learningtransformer-circuits.pub
- Open Source Replication & Commentary on Anthropic's Dictionary Learning Paper — LessWronglesswrong.com
- Matryoshka Sparse Autoencoders — LessWronglesswrong.com
- A gentle introduction to sparse autoencodersnickjiang.substack.com
- Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnettransformer-circuits.pub
- Toy Models of Superpositiontransformer-circuits.pub
- Interpretability with Sparse Autoencoders (Colab exercises) — LessWronglesswrong.com