flâneur — a map of the web's best reading

Vision in Action: Learning Active Perception from Human Demonstrations

vision-in-action.github.io · 1,792 words · saved by 1 readers

Abstract: We present Vision in Action (ViA), an active perception system for bimanual robot manipulation. ViA learns task-relevant active perceptual strategies (e.g., searching, tracking, and focusing) directly from human demonstrations. On the hardware side, ViA employs a simple yet effective 6-DoF robotic neck to enable flexible, human-like head movements. To capture human active perception strategies, we design a VR-based teleoperation interface that creates a shared observation space between the robot and the human operator. To mitigate VR motion sickness caused by latency in the robot’s physical movements, the interface uses an intermediate 3D scene representation, enabling real-time view rendering on the operator side while asynchronously updating the scene with the robot’s latest observations. Together, these design elements enable the learning of robust visuomotor policies for three complex, multi-stage bimanual manipulation tasks involving visual occlusions, significantly outp

Vision in Action: Learning Active Perception from Human Demonstrations --> Vision in Action Learning Active Perception from Human Demonstrations Haoyu Xiong &nbsp&nbsp Xiaomeng Xu &nbsp&nbsp Jimmy Wu &nbsp&nbsp Yifan Hou &nbsp&nbsp Jeannette Bohg &nbsp&nbsp Shuran Song --> --> --> CoRL 2025 Paper Video Code Hardware TL;DR --> The Vision in Action (ViA) system uses a single active head camera for policy learning. --> --> --> We present ViA, an active perception system for bimanual manipulation. --> --> active head camera for policy learning. --> --> Argos Panoptes is a many-eyed giant in Greek

Explore this link on the map →

related reading