Sparse Autoencoders Reveal Interpretable and Steerable Features in VLA Models
Vision-Language-Action (VLA) models have emerged as a promising approach for general-purpose robot manipulation. However, their generalization is inconsistent: while these models can perform impressively in some settings, fine-tuned variants often fail on novel objects, scenes, and instructions. We apply mechanistic interpretability techniques to better understand the inner workings of VLA models. To probe internal representations, we train Sparse Autoencoders (SAEs) on hidden layer activations of the VLA. SAEs learn a sparse dictionary whose features act as a compact, interpretable basis for the model’s computation. We find that the large majority of extracted SAE features correspond to memorized sequences from specific training demonstrations. However, some features correspond to interpretable, general, and steerable motion primitives and semantic properties, offering a promising glimpse toward VLA generalizability. We propose a metric to categorize features according to whether they
Sparse Autoencoders Reveal Interpretable and Steerable Features in VLA Models Aiden Swann 1 , Lachlain McGranahan 3 , Hugo Buurmeijer 3 Monroe Kennedy III 1,2 , Mac Schwager 3 1 Department of Mechanical Engineering, 2 Department of Computer Science 3 Department of Aeronautics & Astronautics Stanford University Abstract Vision-Language-Action (VLA) models have emerged as a promising approach for general-purpose robot manipulation. However, their generalization is inconsistent: while these models can perform impressively in some settings, fine-tuned variants often fail on novel objects, scenes,
Explore this link on the map →saved by
related reading
- 45d74e190008c7bff2845ffc8e3facd3-Paper-Conference.pdfproceedings.iclr.cc
- [2602.10556] LAP: Language-Action Pre-Training Enables Zero-shot Cross-Embodiment Transferarxiv.org
- e5b5c402bb7bd5e60bede6961d6fe39e-Paper-Conference.pdfproceedings.iclr.cc
- Towards Monosemanticity: Decomposing Language Models With Dictionary Learningtransformer-circuits.pub
- An Intuitive Explanation of Sparse Autoencoders for LLM Interpretability | Adam Karvonenadamkarvonen.github.io
- Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnettransformer-circuits.pub
- State of Vision-Language-Action (VLA) Research at ICLR 2026 – Moritz Reussmbreuss.github.io
- A VLA with Open-World Generalizationpi.website
- A gentle introduction to sparse autoencoders — LessWronglesswrong.com
- pdfopenreview.net
- Sparse Autoencoders Reveal Universal Feature Spaces Across Large Language Modelsarxiv.org
- Sparse Autoencoders Find Highly Interpretable Features in Language Modelsarxiv.org