The Bottlenecks to Scaling Foundation Models for Robotics | ICLR Blogposts 2026
Current approaches to building Vision-Language-Action (VLA) models largely rely on combining pre-trained Vision-Language Models (VLMs) with imitation learning. While effective in narrow benchmarks, this paradigm faces fundamental limitations for developing general-purpose robots that operate in complex, dynamic environments. In this article, I first review the standard training recipe and identify key bottlenecks, drawing on both my observations and existing empirical evidence. I then outline a path forward: integrating online reinforcement learning with pre-trained VLMs to enable lightweight, computationally efficient methods that scale with available resources. Anonymous Anonymous April 27, 2026 Will Scaling Solve Robotics? This question has been circulating widely in the robotics community in the last couple of years . When I first encountered it, I found it oddly vague. Scaling what - data, compute, memory, neural network capacity? And for which algorithms and settings? As I probed
Explore this link on the map →