flâneur — a map of the web's best reading

TVL: A Touch, Vision, and Language Dataset for Multimodal Alignment

tactile-vlm.github.io · 969 words · saved by 1 readers

We introduce the Touch-Vision-Language (TVL) dataset, which combines paired tactile and visual observations with both human-annotated and VLM-generated tactile-semantic labels. We then leverage a contrastive learning approach to train a CLIP-aligned tactile encoder and finetune an open-source LLM for a tactile description task. Our results show that incorporating tactile information allows us to significantly outperform state-of-the-art VLMs (including the label generating model) on a tactile understanding task. We gather data using a handheld, 3D printed collection device. Tactile data are collected using a DIGIT sensor: a compact, open-source tactile sensor that provides observations in the form of RGB images of a deformable internal surface. Image data come from a Logitech BRIO webcam, positioned such that the tactile sensor and the point of contact are within its field of view. The collected data are then temporally synchronized and labeled with language descriptions of the tactile

TVL: A Touch, Vision, and Language Dataset for Multimodal Alignment TVL A Touch, Vision, and Language Dataset for Multimodal Alignment Max (Letian) Fu 1 Gaurav Datta * 1 Raven Huang * 1 Will Panitch * 1 Jaimyn Drake * 1 Joseph Ortiz 2 Mustafa Mukadam 2 Mike Lambeta 2 Roberto Calandra 3,4 Ken Goldberg 1 1 UC Berkeley 2 Meta AI Research 3 TU Dresden 4 The Centre for Tactile Internet with Human-in-the-Loop (CeTI) Paper Code Dataset Models Citation TL;DR : Multi-modal alignment made easy using GPT-4V pseudolabels. Overview We introduce the Touch-Vision-Language (TVL) dataset, which combines paired

Explore this link on the map →

related reading