Llama Scope: Extracting Features from Llama 3.1-8B with SAEs
arxiv.org · 4,115 words · saved by 1 readers
N/A
Llama Scope: Extracting Millions of Features from Llama-3.1-8B with Sparse Autoencoders L LAMA S COPE : E XTRACTING M ILLIONS OF F EA - TURES FROM L LAMA -3.1-8B WITH S PARSE AUTOEN - CODERS Zhengfu He1,2 Wentao Shu1 Xuyang Ge1 Lingjie Chen1 Junxuan Wang1 Yunhua Zhou1,3 Frances Liu Qipeng Guo1,2,3 Xuanjing Huang1 Zuxuan Wu1…
related reading
- An Intuitive Explanation of Sparse Autoencoders for LLM Interpretability | Adam Karvonenadamkarvonen.github.io
- A gentle introduction to sparse autoencoders — LessWronglesswrong.com
- A Survey on Sparse Autoencoders: Interpreting the Internal Mechanisms of Large Language Modelsarxiv.org
- Towards Monosemanticity: Decomposing Language Models With Dictionary Learningtransformer-circuits.pub
- Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnettransformer-circuits.pub
- pdfopenreview.net
- A gentle introduction to sparse autoencodersnickjiang.substack.com
- GitHub - PaulPauls/llama3_interpretability_sae: A complete end-to-end pipeline for LLM interpretability with sparse autoencoders (SAEs) using Llama 3.2, written in pure PyTorch and fully reproducible.github.com
- An Intuitive Explanation of Sparse Autoencoders for Mechanistic Interpretability of LLMs — LessWronglesswrong.com
- Towards Monosemanticity: Decomposing Language Models With Dictionary Learningtransformer-circuits.pub
- Matryoshka Sparse Autoencoders — LessWronglesswrong.com
- Sparse Autoencoders Find Highly Interpretable Features in Language Modelsarxiv.org