Interpretability with Sparse Autoencoders (Colab exercises) — LessWrong
Update (13th October 2024) - these exercises have been significantly expanded on. Now there are 2 exercise sets: the first one dives deeply into theo…
x Interpretability with Sparse Autoencoders (Colab exercises) — LessWrong Sparse Autoencoders (SAEs) Exercises / Problem-Sets Interpretability (ML & AI) Superposition AI Frontpage 83 Interpretability with Sparse Autoencoders (Colab exercises) by CallumMcDougall 29th Nov 2023 AI Alignment Forum 4 min read 9 83 Ω 32 Update (13th October 2024) - these exercises have been significantly expanded on. Now there are 2 exercise sets: the first one dives deeply into theoretical topics related to superposition, while the second one (much larger) includes a streamlined version of the first one, as well as
Explore this link on the map →related reading
- An Intuitive Explanation of Sparse Autoencoders for LLM Interpretability | Adam Karvonenadamkarvonen.github.io
- A gentle introduction to sparse autoencoders — LessWronglesswrong.com
- Towards Monosemanticity: Decomposing Language Models With Dictionary Learningtransformer-circuits.pub
- An Intuitive Explanation of Sparse Autoencoders for Mechanistic Interpretability of LLMs — LessWronglesswrong.com
- A gentle introduction to sparse autoencodersnickjiang.substack.com
- pdfopenreview.net
- Toy Models of Superpositiontransformer-circuits.pub
- Sparse Autoencoders Find Highly Interpretable Features in Language Modelsarxiv.org
- Open Source Replication & Commentary on Anthropic's Dictionary Learning Paper — LessWronglesswrong.com
- Matryoshka Sparse Autoencoders — LessWronglesswrong.com
- Do sparse autoencoders find "true features"? — LessWronglesswrong.com
- Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnettransformer-circuits.pub