SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model Interpretability
arxiv.org · 6,315 words · saved by 1 readers
N/A
SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model Interpretability Adam Karvonen * 1 Can Rager * 1 Johnny Lin * 2 Curt Tigges * 2 Joseph Bloom * 2 David Chanin 3 Yeu-Tong Lau 1 Eoin Farrell 1 Callum McDougall Kola Ayonrinde Demian Till 4 Matthew Wearden 5 Arthur Conmy Samuel Marks 6 Neel Nanda…
related reading
- An Intuitive Explanation of Sparse Autoencoders for LLM Interpretability | Adam Karvonenadamkarvonen.github.io
- A Survey on Sparse Autoencoders: Interpreting the Internal Mechanisms of Large Language Modelsarxiv.org
- Towards Monosemanticity: Decomposing Language Models With Dictionary Learningtransformer-circuits.pub
- A gentle introduction to sparse autoencoders — LessWronglesswrong.com
- Sparse Autoencoders Find Highly Interpretable Features in Language Modelsarxiv.org
- An Intuitive Explanation of Sparse Autoencoders for Mechanistic Interpretability of LLMs — LessWronglesswrong.com
- pdfopenreview.net
- Towards Monosemanticity: Decomposing Language Models With Dictionary Learningtransformer-circuits.pub
- A gentle introduction to sparse autoencodersnickjiang.substack.com
- Llama Scope: Extracting Features from Llama 3.1-8B with SAEsarxiv.org
- [2409.14507] A is for Absorption: Studying Feature Splitting and Absorption in Sparse Autoencodersar5iv.labs.arxiv.org
- I Trained a Language Model. Then I Built a Brain Scanner and Looked Inside It. | by Caleb DeLeeuw | Mediummedium.com