flâneur — a map of the web's best reading

Multimodal interpretability in 2024

soniajoseph.ai · 4,412 words · saved by 1 readers

Multimodal interpretability, from sparse feature circuits with SAEs, to vision transformers leveraging CLIP's shared text-image space.

I'm writing this post to clarify my thoughts and update my collaborators on multimodal interpretability in 2024. Having spent part of the summer in the AI safety sphere in Berkeley, and then joining the video understanding team at FAIR as a visiting researcher, I'm bridging two communities: the language mechanistic interpretability efforts in AI safety, and the efficiency-focused Vision-Language Model (VLM) community in industry. Some content may be more familiar to one community than the other. As part of a broader series, this post is a progress update on my thinking around multimodal interp

Explore this link on the map →

saved by

related reading