[2602.06001] Visuo-Tactile World Models
Abstract:We introduce multi-task Visuo-Tactile World Models (VT-WM), which capture the physics of contact through touch reasoning. By complementing vision with tactile sensing, VT-WM better understands robot-object interactions in contact-rich tasks, avoiding common failure modes of vision-only models under occlusion or ambiguous contact states, such as objects disappearing, teleporting, or moving in ways that violate basic physics. Trained across a set of contact-rich manipulation tasks, VT-WM improves physical fidelity in imagination, achieving 33% better performance at maintaining object permanence and 29% better compliance with the laws of motion in autoregressive rollouts. Moreover, experiments show that grounding in contact dynamics also translates to planning. In zero-shot real-robot experiments, VT-WM achieves up to 35% higher success rates, with the largest gains in multi-step, contact-rich tasks. Finally, VT-WM demonstrates significant downstream versatility, effectively adapting its learned contact dynamics to a novel task and achieving reliable planning success with only a limited set of demonstrations.
# link_2pnte9u73d.pdf ## Metadata - PDFFormatVersion=1.7 - IsLinearized=false - IsAcroFormPresent=false - IsXFAPresent=false - IsCollectionPresent=false - IsSignaturesPresent=false - Author=Carolina Higuera; Sergio Arnaud; Byron Boots; Mustafa Mukadam; Francois Robert Hogan; Franziska Meier - Creator=arXiv GenPDF (tex2pdf:57610bf) - Custom.DOI=https://doi.org/10.48550/arXiv.2602.06001 - Custom.License=http://arxiv.org/licenses/nonexclusive-distrib/1.0/ - Custom.PTEX.Fullbanner=This is pdfTeX, Version 3.141592653-2.6-1.40.28 (TeX Live 2025) kpathsea version 6.4.1 - Custom.arXivID=https://arxiv.
saved by
related reading
- OSMO: Open-Source Tactile Glove for Human-to-Robot Skill Transferarxiv.org
- The flavor of the bitter lesson for computer vision - Vincent Sitzmannvincentsitzmann.com
- World Models | Rohit Bandarurohitbandaru.github.io
- A Functional Taxonomy of World Models - Dr. Fei-Fei Lidrfeifei.substack.com
- How Claude Performs on Robotics Tasks \ Anthropicanthropic.com
- LeWorldModel: Stable End-to-End Joint-Embedding Predictive Architecture from Pixelsle-wm.github.io
- [2509.02722] Planning with Reasoning using Vision Language World Modelarxiv.org
- Explore | alphaXivalphaxiv.org
- Anirudha Majumdar (@Majumdar_Ani) on Xx.com
- 1X World Model | From Video to Action: A New Way Robots Learn1x.tech
- Robot Synesthesia: In-Hand Manipulation with Visuotactile Sensingarxiv.org
- [2606.19161] HT-Bench: Benchmarking and Learning Dexterous Full-Hand Tactile Representations with Egocentric Visionarxiv.org