flâneur — a map of the web's best reading

Sparse Autoencoders Reveal Interpretable and Steerable Features in VLA Models

arxiv.org · 12,657 words · saved by 1 readers

Vision-Language-Action (VLA) models have emerged as a promising approach for general-purpose robot manipulation. However, their generalization is inconsistent: while these models can perform impressively in some settings, fine-tuned variants often fail on novel objects, scenes, and instructions. We apply mechanistic interpretability techniques to better understand the inner workings of VLA models. To probe internal representations, we train Sparse Autoencoders (SAEs) on hidden layer activations of the VLA. SAEs learn a sparse dictionary whose features act as a compact, interpretable basis for the model’s computation. We find that the large majority of extracted SAE features correspond to memorized sequences from specific training demonstrations. However, some features correspond to interpretable, general, and steerable motion primitives and semantic properties, offering a promising glimpse toward VLA generalizability. We propose a metric to categorize features according to whether they

Sparse Autoencoders Reveal Interpretable and Steerable Features in VLA Models Aiden Swann 1 , Lachlain McGranahan 3 , Hugo Buurmeijer 3 Monroe Kennedy III 1,2 , Mac Schwager 3 1 Department of Mechanical Engineering, 2 Department of Computer Science 3 Department of Aeronautics & Astronautics Stanford University Abstract Vision-Language-Action (VLA) models have emerged as a promising approach for general-purpose robot manipulation. However, their generalization is inconsistent: while these models can perform impressively in some settings, fine-tuned variants often fail on novel objects, scenes,

Explore this link on the map →

saved by

related reading