Multimodal interpretability in 2024
soniajoseph.ai · 4,412 words · saved by 1 readers
Multimodal interpretability, from sparse feature circuits with SAEs, to vision transformers leveraging CLIP's shared text-image space.
I'm writing this post to clarify my thoughts and update my collaborators on multimodal interpretability in 2024. Having spent part of the summer in the AI safety sphere in Berkeley, and then joining the video understanding team at FAIR as a visiting researcher, I'm bridging two communities: the language mechanistic interpretability efforts in AI safety, and the efficiency-focused Vision-Language Model (VLM) community in industry. Some content may be more familiar to one community than the other. As part of a broader series, this post is a progress update on my thinking around multimodal interp
saved by
related reading
- Transformer Circuits Threadtransformer-circuits.pub
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- Towards Monosemanticity: Decomposing Language Models With Dictionary Learningtransformer-circuits.pub
- Prism: mapping interpretable concepts and features in a latent space of language | thesephist.comthesephist.com
- Towards Multimodal Interpretability: Learning Sparse Interpretable Features in Vision Transformers — LessWronglesswrong.com
- Dario Amodei — The Urgency of Interpretabilitydarioamodei.com
- Tracing the thoughts of a large language model \ Anthropicanthropic.com
- Circuit Tracing in Vision–Language Models:Understanding the Internal Mechanisms of Multimodal Thinkingarxiv.org
- Towards Monosemanticity: Decomposing Language Models With Dictionary Learningtransformer-circuits.pub
- The Building Blocks of Interpretabilitydistill.pub
- A Comprehensive Mechanistic Interpretability Explainer & Glossary — Neel Nandaneelnanda.io