flâneur — a map of the web's best reading

ASIMOV Benchmark v1

asimov-benchmark.github.io · 446 words · saved by 1 readers

Until recently, robotics safety research was predominantly about collision avoidance and hazard reduction in the immediate vicinity of a robot. Since the advent of large vision and language models (VLMs), robots are now also capable of higher-level semantic scene understanding and natural language interactions with humans. Despite their known vulnerabilities (e.g. hallucinations or jail-breaking), VLMs are being handed control of robots capable of physical contact with the real world. This can lead to dangerous behaviors, making semantic safety for robots a matter of immediate concern. Our contributions in this paper are two fold: first, to address these emerging risks, we release the ASIMOV Benchmark — a large-scale and comprehensive collection of datasets for evaluating and improving semantic safety of foundation models serving as robot brains. Our data generation recipe is highly scalable: by leveraging text and image generation techniques, we generate undesirable situations from re

ASIMOV Benchmark v1 ASIMOV Benchmark v1 Generating Robot Constitutions & Benchmarks for Semantic Safety Pierre Sermanet 1 , Anirudha Majumdar 1,2 , Alex Irpan 1 , Dmitry Kalashnikov 1 , Vikas Sindhwani 1 , 1 Google DeepMind, 2 Princeton University Paper Cite arXiv Data Code Video Accepted at CoRL 2025 Other versions: v2 Your browser does not support the video tag. Abstract Until recently, robotics safety research was predominantly about collision avoidance and hazard reduction in the immediate vicinity of a robot. Since the advent of large vision and language models (VLMs), robots are now also

Explore this link on the map →

related reading