A gentle introduction to sparse autoencoders — LessWrong
Sparse autoencoders (SAEs) are the current hot topic 🔥 in the interpretability world. In late May, Anthropic released a paper that shows how to use sparse autoencoders to effectively break down the internal reasoning of Claude 3 (Anthropic’s LLM) . Shortly after, OpenAI published a paper successfully applying a similar procedure for GPT4. What’s exciting about SAEs is that, for the first time, they provide a scalable method to peer inside virtually any transformer-based LLM today. In this introduction, I want to unpack the jargon, intuition, and technical details behind SAEs. This post is targeted for technical folks without a background in interpretability; it extracts what I believe are the most important insights in the history leading up to SAEs. I’ll mainly focus on the Anthropic line of research, but there are many research groups, academic and industry, that have played pivotal roles in getting to where we are[1]. Since the moment LLMs took off in the early 2020s, we’ve really
x A gentle introduction to sparse autoencoders — LessWrong Sparse Autoencoders (SAEs) AI Personal Blog 25 A gentle introduction to sparse autoencoders by Nick Jiang 2nd Sep 2024 8 min read 2 25 Sparse autoencoders (SAEs) are the current hot topic 🔥 in the interpretability world. In late May, Anthropic released a paper that shows how to use sparse autoencoders to effectively break down the internal reasoning of Claude 3 (Anthropic’s LLM) . Shortly after, OpenAI published a paper successfully applying a similar procedure for GPT4. What’s exciting about SAEs is that, for the first time, they pro
Explore this link on the map →saved by
related reading
- An Intuitive Explanation of Sparse Autoencoders for LLM Interpretability | Adam Karvonenadamkarvonen.github.io
- A gentle introduction to sparse autoencodersnickjiang.substack.com
- An Intuitive Explanation of Sparse Autoencoders for Mechanistic Interpretability of LLMs — LessWronglesswrong.com
- Towards Monosemanticity: Decomposing Language Models With Dictionary Learningtransformer-circuits.pub
- pdfopenreview.net
- Sparse Autoencoders Find Highly Interpretable Features in Language Modelsarxiv.org
- Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnettransformer-circuits.pub
- [2409.14507] A is for Absorption: Studying Feature Splitting and Absorption in Sparse Autoencodersar5iv.labs.arxiv.org
- Interpretability with Sparse Autoencoders (Colab exercises) — LessWronglesswrong.com
- Do sparse autoencoders find "true features"? — LessWronglesswrong.com
- Open Source Replication & Commentary on Anthropic's Dictionary Learning Paper — LessWronglesswrong.com
- A List of 45+ Mech Interp Project Ideas from Apollo Research’s Interpretability Team — AI Alignment Forumalignmentforum.org