FAST: Efficient Action Tokenization for Vision-Language-Action Models
This is experimental HTML to improve accessibility. We invite you to report rendering errors. Use Alt+Y to toggle on accessible reporting links and Alt+Shift+Y to toggle off. Learn more about this project and help improve conversions. HTML conversions sometimes display errors due to content that did not convert correctly from the source. This paper uses the following packages that are not yet supported by the HTML conversion tool. Feedback on these issues are not necessary; they are known and are being worked on. Authors: achieve the best HTML results from your LaTeX submissions by following these best practices. Autoregressive sequence models, such as Transformer-based vision-language action (VLA) policies, can be tremendously effective for capturing complex and generalizable robotic behaviors. However, such models require us to choose a tokenization of our continuous action signals, which determines how the discrete symbols predicted by the model map to continuous robot actions. We
FAST: Efficient Action Tokenization for Vision-Language-Action Models Karl Pertsch ∗,1,2,3 , Kyle Stachowicz ∗,2 , Brian Ichter 1 , Danny Driess 1 , Suraj Nair 1 , Quan Vuong 1 , Oier Mees 2 , Chelsea Finn 1,3 , Sergey Levine 1,2 1 Physical Intelligence, 2 UC Berkeley, 3 Stanford https://pi.website/research/fast ∗ : Core contributors Correspondence to: research@physicalintelligence.company Abstract Autoregressive sequence models, such as Transformer-based vision-language action (VLA) policies, can be tremendously effective for capturing complex and generalizable robotic behaviors. However, suc
Explore this link on the map →related reading
- The First Fully General Computer Action Model | blogsi.inc
- State of Vision-Language-Action (VLA) Research at ICLR 2026 – Moritz Reussmbreuss.github.io
- 45d74e190008c7bff2845ffc8e3facd3-Paper-Conference.pdfproceedings.iclr.cc
- Precise Manipulation with Efficient Online RLpi.website
- The Bitter Lesson is coming for Tokenization – ⛰️ lucalplucalp.dev
- [2603.05438] Planning in 8 Tokens: A Compact Discrete Tokenizer for Latent World Modelarxiv.org
- 𝜋₀: A Vision-Language-Action Flow Model for General Robot Controlarxiv.org
- Efficient World Models with Context-Aware Tokenizationarxiv.org
- MolmoAct Action Reasoning Models that can Reason in Spacearxiv.org
- FASTERinnovator-zero.github.io
- e5b5c402bb7bd5e60bede6961d6fe39e-Paper-Conference.pdfproceedings.iclr.cc
- Causal Video Models Are Data-Efficient Robot Policy Learners | Rhoda AIrhoda.ai