Towards Multimodal Interpretability: Learning Sparse Interpretable Features in Vision Transformers — LessWrong
Two Minute Summary In this post I present my results from training a Sparse Autoencoder (SAE) on a CLIP Vision Transformer (ViT) using the ImageNet-1…
x Towards Multimodal Interpretability: Learning Sparse Interpretable Features in Vision Transformers — LessWrong Interpretability (ML & AI) Sparse Autoencoders (SAEs) MATS Program AI Frontpage 94 Towards Multimodal Interpretability: Learning Sparse Interpretable Features in Vision Transformers by hugofry 29th Apr 2024 14 min read 9 94 Executive Summary In this post I present my results from training a Sparse Autoencoder (SAE) on a CLIP Vision Transformer (ViT) using the ImageNet-1k dataset. I have created an interactive web app, 'SAE Explorer', to allow the public to explore the visual feature
Explore this link on the map →related reading
- Towards Monosemanticity: Decomposing Language Models With Dictionary Learningtransformer-circuits.pub
- Multimodal interpretability in 2024soniajoseph.ai
- Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnettransformer-circuits.pub
- An Intuitive Explanation of Sparse Autoencoders for LLM Interpretability | Adam Karvonenadamkarvonen.github.io
- A gentle introduction to sparse autoencoders — LessWronglesswrong.com
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnettransformer-circuits.pub
- Chapter 1: Transformer Interpretability - ARENAlearn.arena.education
- Chapter 1: Transformer Interpretability - ARENAlearn.arena.education
- Transformer Circuits Threadtransformer-circuits.pub
- Identifying Functionally Important Features with End-to-End Sparse Dictionary Learning — LessWronglesswrong.com
- A List of 45+ Mech Interp Project Ideas from Apollo Research’s Interpretability Team — AI Alignment Forumalignmentforum.org