Vision in Action: Learning Active Perception from Human Demonstrations
Abstract: We present Vision in Action (ViA), an active perception system for bimanual robot manipulation. ViA learns task-relevant active perceptual strategies (e.g., searching, tracking, and focusing) directly from human demonstrations. On the hardware side, ViA employs a simple yet effective 6-DoF robotic neck to enable flexible, human-like head movements. To capture human active perception strategies, we design a VR-based teleoperation interface that creates a shared observation space between the robot and the human operator. To mitigate VR motion sickness caused by latency in the robot’s physical movements, the interface uses an intermediate 3D scene representation, enabling real-time view rendering on the operator side while asynchronously updating the scene with the robot’s latest observations. Together, these design elements enable the learning of robust visuomotor policies for three complex, multi-stage bimanual manipulation tasks involving visual occlusions, significantly outp
Vision in Action: Learning Active Perception from Human Demonstrations --> Vision in Action Learning Active Perception from Human Demonstrations Haoyu Xiong    Xiaomeng Xu    Jimmy Wu    Yifan Hou    Jeannette Bohg    Shuran Song --> --> --> CoRL 2025 Paper Video Code Hardware TL;DR --> The Vision in Action (ViA) system uses a single active head camera for policy learning. --> --> --> We present ViA, an active perception system for bimanual manipulation. --> --> active head camera for policy learning. --> --> Argos Panoptes is a many-eyed giant in Greek
Explore this link on the map →related reading
- The flavor of the bitter lesson for computer vision - Vincent Sitzmannvincentsitzmann.com
- [2304.13705] Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardwarear5iv.labs.arxiv.org
- Helix: A Vision-Language-Action Model for Generalist Humanoid Controlfigure.ai
- When Models Manipulate Manifolds: The Geometry of a Counting Tasktransformer-circuits.pub
- Precise Manipulation with Efficient Online RLpi.website
- e5b5c402bb7bd5e60bede6961d6fe39e-Paper-Conference.pdfproceedings.iclr.cc
- Homepalanc.github.io
- Building a robotics research setup that lives next to my desk – dfdx labsdfdxlabs.com
- A VLA with Open-World Generalizationpi.website
- OSMO: Open-Source Tactile Glove for Human-to-Robot Skill Transferarxiv.org
- SARM2: Multi-Task Stage Aware Reward Modeling for Self Improving Robotic Manipulationqianzhong-chen.github.io
- Lift, Splat, Shoot: Encoding Images from Arbitrary Camera Rigs by Implicitly Unprojecting to 3Dresearch.nvidia.com