✳flâneur — a map of the web's best reading
openreview.net · 19,874 words · saved by 1 readers
N/A
# link_1nkupw53n51.pdf ## Metadata - PDFFormatVersion=1.5 - IsLinearized=true - IsAcroFormPresent=false - IsXFAPresent=false - IsCollectionPresent=false - IsSignaturesPresent=false - CreationDate=D:20251022160513Z - Creator=LaTeX with hyperref - ModDate=D:20251022160513Z - Custom.PTEX.Fullbanner=This is pdfTeX, Version 3.141592653-2.6-1.40.26 (TeX Live 2024) kpathsea version 6.4.0 - Producer=pdfTeX-1.40.26 - Trapped=False ## Contents ### Page 1 A is for Absorption: Studying Feature Splitting and Absorption in Sparse AutoencodersDavid Chanin*,1,2 James Wilken-Smith*,1 Tomáš Dulka*,1 Hardik B
Explore this link on the map →saved by
related reading
- [2409.14507] A is for Absorption: Studying Feature Splitting and Absorption in Sparse Autoencodersar5iv.labs.arxiv.org
- An Intuitive Explanation of Sparse Autoencoders for LLM Interpretability | Adam Karvonenadamkarvonen.github.io
- Towards Monosemanticity: Decomposing Language Models With Dictionary Learningtransformer-circuits.pub
- Matryoshka Sparse Autoencoders — LessWronglesswrong.com
- A gentle introduction to sparse autoencoders — LessWronglesswrong.com
- Sparse Autoencoders Find Highly Interpretable Features in Language Modelsarxiv.org
- A gentle introduction to sparse autoencodersnickjiang.substack.com
- Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnettransformer-circuits.pub
- An Intuitive Explanation of Sparse Autoencoders for Mechanistic Interpretability of LLMs — LessWronglesswrong.com
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- Identifying Functionally Important Features with End-to-End Sparse Dictionary Learning — LessWronglesswrong.com
- Do sparse autoencoders find "true features"? — LessWronglesswrong.com