TVL: A Touch, Vision, and Language Dataset for Multimodal Alignment
We introduce the Touch-Vision-Language (TVL) dataset, which combines paired tactile and visual observations with both human-annotated and VLM-generated tactile-semantic labels. We then leverage a contrastive learning approach to train a CLIP-aligned tactile encoder and finetune an open-source LLM for a tactile description task. Our results show that incorporating tactile information allows us to significantly outperform state-of-the-art VLMs (including the label generating model) on a tactile understanding task. We gather data using a handheld, 3D printed collection device. Tactile data are collected using a DIGIT sensor: a compact, open-source tactile sensor that provides observations in the form of RGB images of a deformable internal surface. Image data come from a Logitech BRIO webcam, positioned such that the tactile sensor and the point of contact are within its field of view. The collected data are then temporally synchronized and labeled with language descriptions of the tactile
TVL: A Touch, Vision, and Language Dataset for Multimodal Alignment TVL A Touch, Vision, and Language Dataset for Multimodal Alignment Max (Letian) Fu 1 Gaurav Datta * 1 Raven Huang * 1 Will Panitch * 1 Jaimyn Drake * 1 Joseph Ortiz 2 Mustafa Mukadam 2 Mike Lambeta 2 Roberto Calandra 3,4 Ken Goldberg 1 1 UC Berkeley 2 Meta AI Research 3 TU Dresden 4 The Centre for Tactile Internet with Human-in-the-Loop (CeTI) Paper Code Dataset Models Citation TL;DR : Multi-modal alignment made easy using GPT-4V pseudolabels. Overview We introduce the Touch-Vision-Language (TVL) dataset, which combines paired
Explore this link on the map →related reading
- Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Modelsarxiv.org
- State of Vision-Language-Action (VLA) Research at ICLR 2026 – Moritz Reussmbreuss.github.io
- cs.unc.edu/~mbansal/teaching/nlp-comp790-590-spring23.htmlcs.unc.edu
- 2403.09611.pdfarxiv.org
- GitHub - jacobhilton/deep_learning_curriculum: Language model alignment-focused deep learning curriculum · GitHubgithub.com
- Subliminal Learning: Language Models Transmit Behavioral Traits via Hidden Signals in Dataalignment.anthropic.com
- The Platonic Representation Hypothesisphillipi.github.io
- Emergence of Human to Robot Transfer in Vision-Language-Action Modelspi.website
- Unified Multimodal Models as Auto-Encodersarxiv.org
- [2602.06001] Visuo-Tactile World Modelsarxiv.org
- A VLA with Open-World Generalizationpi.website
- Chunyuan Lichunyuan.li