Identifying Functionally Important Features with End-to-End Sparse Dictionary Learning — LessWrong
lesswrong.com · 4,305 words · saved by 1 readers
A short summary of the paper is presented below. …
x Identifying Functionally Important Features with End-to-End Sparse Dictionary Learning — LessWrong Apollo Research (org) Interpretability (ML & AI) Sparse Autoencoders (SAEs) AI Frontpage 57 Identifying Functionally Important Features with End-to-End Sparse Dictionary Learning by Dan Braun , Jordan Taylor , Nicholas Goldowsky-Dill , Lee Sharkey 17th May 2024 AI Alignment Forum Linkpost for arxiv.org 5 min read 20 57 Ω 28 A short summary of the paper is presented below. This work was produced by Apollo Research in collaboration with Jordan Taylor (MATS + University of Queensland) . TL;DR: We
saved by
related reading
- Towards Monosemanticity: Decomposing Language Models With Dictionary Learningtransformer-circuits.pub
- An Intuitive Explanation of Sparse Autoencoders for LLM Interpretability | Adam Karvonenadamkarvonen.github.io
- Sparse Autoencoders Find Highly Interpretable Features in Language Modelsarxiv.org
- Towards Monosemanticity: Decomposing Language Models With Dictionary Learningtransformer-circuits.pub
- A gentle introduction to sparse autoencoders — LessWronglesswrong.com
- pdfopenreview.net
- Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnettransformer-circuits.pub
- Matryoshka Sparse Autoencoders — LessWronglesswrong.com
- Open Source Replication & Commentary on Anthropic's Dictionary Learning Paper — LessWronglesswrong.com
- Efficient Dictionary Learning with Switch Sparse Autoencoders — LessWronglesswrong.com
- An Introduction to Exemplar Partitioning for Mechanistic Interpretability — LessWronglesswrong.com
- Improving Dictionary Learning with Gated Sparse Autoencodersarxiv.org