A Survey on Sparse Autoencoders: Interpreting the Internal Mechanisms of Large Language Models
arxiv.org · 6,328 words · saved by 2 readers
N/A
A Survey on Sparse Autoencoders: Interpreting the Internal Mechanisms of Large Language Models Dong Shu1,† , Xuansheng Wu2,† , Haiyan Zhao3,† , Daking Rai4 , Ziyu Yao4 , Ninghao Liu2 , Mengnan Du3 1 Northwestern University 2 University of Georgia…
saved by
related reading
- An Intuitive Explanation of Sparse Autoencoders for LLM Interpretability | Adam Karvonenadamkarvonen.github.io
- A gentle introduction to sparse autoencoders — LessWronglesswrong.com
- Towards Monosemanticity: Decomposing Language Models With Dictionary Learningtransformer-circuits.pub
- An Intuitive Explanation of Sparse Autoencoders for Mechanistic Interpretability of LLMs — LessWronglesswrong.com
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- A gentle introduction to sparse autoencodersnickjiang.substack.com
- Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnettransformer-circuits.pub
- Towards Monosemanticity: Decomposing Language Models With Dictionary Learningtransformer-circuits.pub
- Llama Scope: Extracting Features from Llama 3.1-8B with SAEsarxiv.org
- pdfopenreview.net
- GitHub - PaulPauls/llama3_interpretability_sae: A complete end-to-end pipeline for LLM interpretability with sparse autoencoders (SAEs) using Llama 3.2, written in pure PyTorch and fully reproducible.github.com
- Sparse Autoencoders Find Highly Interpretable Features in Language Modelsarxiv.org