flâneur — a map of the web's best reading

Towards Multimodal Interpretability: Learning Sparse Interpretable Features in Vision Transformers — LessWrong

lesswrong.com · 5,076 words · saved by 1 readers

Two Minute Summary In this post I present my results from training a Sparse Autoencoder (SAE) on a CLIP Vision Transformer (ViT) using the ImageNet-1…

x Towards Multimodal Interpretability: Learning Sparse Interpretable Features in Vision Transformers — LessWrong Interpretability (ML & AI) Sparse Autoencoders (SAEs) MATS Program AI Frontpage 94 Towards Multimodal Interpretability: Learning Sparse Interpretable Features in Vision Transformers by hugofry 29th Apr 2024 14 min read 9 94 Executive Summary In this post I present my results from training a Sparse Autoencoder (SAE) on a CLIP Vision Transformer (ViT) using the ImageNet-1k dataset. I have created an interactive web app, 'SAE Explorer', to allow the public to explore the visual feature

Explore this link on the map →

related reading