✳flâneur — a map of the web's best reading
A gentle introduction to sparse autoencoders
nickjiang.substack.com · 2,606 words · saved by 1 readers
Our biggest leap so far in deciphering how large language models work
A gentle introduction to sparse autoencoders Our biggest leap so far in deciphering how large language models work Nick Jiang Jul 03, 2024 6 Share DALLE3: a lab technician wielding a sledgehammer to break open a black box Sparse autoencoders (SAEs) are the current hot topic 🔥 in the interpretability world. In late May, Anthropic released a paper that shows how to use sparse autoencoders to effectively break down the internal reasoning of Claude 3 (Anthropic’s LLM) 1 . Shortly after, OpenAI published a paper successfully applying a similar procedure for GPT4. What’s exciting about SAEs is that
Explore this link on the map →related reading
- An Intuitive Explanation of Sparse Autoencoders for LLM Interpretability | Adam Karvonenadamkarvonen.github.io
- A gentle introduction to sparse autoencoders — LessWronglesswrong.com
- Towards Monosemanticity: Decomposing Language Models With Dictionary Learningtransformer-circuits.pub
- An Intuitive Explanation of Sparse Autoencoders for Mechanistic Interpretability of LLMs — LessWronglesswrong.com
- Sparse Autoencoders Find Highly Interpretable Features in Language Modelsarxiv.org
- pdfopenreview.net
- [2409.14507] A is for Absorption: Studying Feature Splitting and Absorption in Sparse Autoencodersar5iv.labs.arxiv.org
- Do sparse autoencoders find "true features"? — LessWronglesswrong.com
- Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnettransformer-circuits.pub
- Open Source Replication & Commentary on Anthropic's Dictionary Learning Paper — LessWronglesswrong.com
- I Trained a Language Model. Then I Built a Brain Scanner and Looked Inside It. | by Caleb DeLeeuw | Mediummedium.com
- Interpretability with Sparse Autoencoders (Colab exercises) — LessWronglesswrong.com