✳flâneur — a map of the web's best reading
Sparse Autoencoders Find Highly Interpretable Features in Language Models
arxiv.org · 7,657 words · saved by 2 readers
N/A
# link_2co6mhcts99.pdf ## Metadata - PDFFormatVersion=1.5 - IsLinearized=false - IsAcroFormPresent=false - IsXFAPresent=false - IsCollectionPresent=false - IsSignaturesPresent=false - CreationDate=D:20231005020919Z - Creator=LaTeX with hyperref - ModDate=D:20231005020919Z - Custom.PTEX.Fullbanner=This is pdfTeX, Version 3.141592653-2.6-1.40.25 (TeX Live 2023) kpathsea version 6.3.5 - Producer=pdfTeX-1.40.25 - Trapped=False ## Contents ### Page 1 SPARSE AUTOENCODERS FIND HIGHLY INTER-PRETABLE FEATURES IN LANGUAGE MODELSHoagy Cunningham∗12, Aidan Ewart∗13, Logan Riggs∗1, Robert Huben, Lee Sha
Explore this link on the map →saved by
related reading
- Towards Monosemanticity: Decomposing Language Models With Dictionary Learningtransformer-circuits.pub
- An Intuitive Explanation of Sparse Autoencoders for LLM Interpretability | Adam Karvonenadamkarvonen.github.io
- Towards Monosemanticity: Decomposing Language Models With Dictionary Learningtransformer-circuits.pub
- Toy Models of Superpositiontransformer-circuits.pub
- [2309.08600] Sparse Autoencoders Find Highly Interpretable Features in Language Modelsarxiv.org
- pdfopenreview.net
- A gentle introduction to sparse autoencoders — LessWronglesswrong.com
- A gentle introduction to sparse autoencodersnickjiang.substack.com
- Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnettransformer-circuits.pub
- Prism: mapping interpretable concepts and features in a latent space of language | thesephist.comthesephist.com
- [Interim research report] Taking features out of superposition with sparse autoencoders — LessWronglesswrong.com
- Open Source Replication & Commentary on Anthropic's Dictionary Learning Paper — LessWronglesswrong.com