Sparse Autoencoders Reveal Interpretable and Steerable Features in VLA Models
Vision-Language-Action (VLA) models have emerged as a promising approach for general-purpose robot manipulation. However, their generalization is inconsistent: while these models can perform impressively in some settings, fine-tuned variants often fail on novel objects, scenes, and instructions. We apply mechanistic interpretability techniques to better understand the inner workings of VLA models. To probe internal representations, we train Sparse Autoencoders (SAEs) on hidden layer activations of the VLA. SAEs learn a sparse dictionary whose features act as a compact, interpretable basis for the model’s computation. We find that the large majority of extracted SAE features correspond to memorized sequences from specific training demonstrations. However, some features correspond to interpretable, general, and steerable motion primitives and semantic properties, offering a promising glimpse toward VLA generalizability. We propose a metric to categorize features according to whether they
Sparse Autoencoders Reveal Interpretable and Steerable Features in VLA Models Aiden Swann 1 , Lachlain McGranahan 3 , Hugo Buurmeijer 3 Monroe Kennedy III 1,2 , Mac Schwager 3 1 Department of Mechanical Engineering, 2 Department of Computer Science 3 Department of Aeronautics & Astronautics Stanford University Abstract Vision-Language-Action (VLA) models have emerged as a promising approach for general-purpose robot manipulation. However, their generalization is inconsistent: while these models can perform impressively in some settings, fine-tuned variants often fail on novel objects, scenes,
saved by
related reading
- 45d74e190008c7bff2845ffc8e3facd3-Paper-Conference.pdfproceedings.iclr.cc
- [2602.10556] LAP: Language-Action Pre-Training Enables Zero-shot Cross-Embodiment Transferarxiv.org
- An Intuitive Explanation of Sparse Autoencoders for LLM Interpretability | Adam Karvonenadamkarvonen.github.io
- Towards Monosemanticity: Decomposing Language Models With Dictionary Learningtransformer-circuits.pub
- State of Vision-Language-Action (VLA) Research at ICLR 2026 – Moritz Reussmbreuss.github.io
- A gentle introduction to sparse autoencoders — LessWronglesswrong.com
- Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnettransformer-circuits.pub
- SimpleVLA-RL: Scaling VLA Training via Reinforcement Learningalphaxiv.org
- Llama Scope: Extracting Features from Llama 3.1-8B with SAEsarxiv.org
- A VLA with Open-World Generalizationpi.website
- Towards Monosemanticity: Decomposing Language Models With Dictionary Learningtransformer-circuits.pub
- A Survey on Sparse Autoencoders: Interpreting the Internal Mechanisms of Large Language Modelsarxiv.org