✳flâneur — a map of the web's best reading
Identifying Functionally Important Features with End-to-End Sparse Dictionary Learning — LessWrong
lesswrong.com · 4,305 words · saved by 1 readers
A short summary of the paper is presented below. …
x Identifying Functionally Important Features with End-to-End Sparse Dictionary Learning — LessWrong Apollo Research (org) Interpretability (ML & AI) Sparse Autoencoders (SAEs) AI Frontpage 57 Identifying Functionally Important Features with End-to-End Sparse Dictionary Learning by Dan Braun , Jordan Taylor , Nicholas Goldowsky-Dill , Lee Sharkey 17th May 2024 AI Alignment Forum Linkpost for arxiv.org 5 min read 20 57 Ω 28 A short summary of the paper is presented below. This work was produced by Apollo Research in collaboration with Jordan Taylor (MATS + University of Queensland) . TL;DR: We
Explore this link on the map →saved by
related reading
- Towards Monosemanticity: Decomposing Language Models With Dictionary Learningtransformer-circuits.pub
- An Intuitive Explanation of Sparse Autoencoders for LLM Interpretability | Adam Karvonenadamkarvonen.github.io
- Sparse Autoencoders Find Highly Interpretable Features in Language Modelsarxiv.org
- pdfopenreview.net
- Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnettransformer-circuits.pub
- A gentle introduction to sparse autoencoders — LessWronglesswrong.com
- Matryoshka Sparse Autoencoders — LessWronglesswrong.com
- Open Source Replication & Commentary on Anthropic's Dictionary Learning Paper — LessWronglesswrong.com
- Efficient Dictionary Learning with Switch Sparse Autoencoders — LessWronglesswrong.com
- Improving Dictionary Learning with Gated Sparse Autoencodersarxiv.org
- A List of 45+ Mech Interp Project Ideas from Apollo Research’s Interpretability Team — AI Alignment Forumalignmentforum.org
- Do sparse autoencoders find "true features"? — LessWronglesswrong.com