flâneur — a map of the web's best reading

Anirudha Majumdar on X: "Should we predict pixels for world models in robotics?" / X

x.com · saved by 1 readers

To view keyboard shortcuts, press question mark View keyboard shortcuts Home Explore Notifications 1 Chat Grok Premium Bookmarks Creator Studio Articles Profile More Post shiza @ShizaCharania Article See new posts Conversation Anirudha Majumdar @Majumdar_Ani Should we predict pixels for world models in robotics? 8 58 344 39K World models for robotics Within the robotics community, there seems to be a general consensus floating in the air: generalist policies of the future will be built on a “world modeling” recipe rather than the VLM-backbone approach that has dominated thus far. The argument goes as follows. VLMs are not explicitly trained to predict the future, and are thus unreliable at the kind of geometric, spatial, and physical reasoning abilities needed to predict fine-grained consequences of actions. In contrast, a world model allows a robot to “imagine” the future in order to plan, e.g., by (1) generating a video of imagined success and using an inverse dynamics model to infer

Explore this link on the map →

saved by