ROJun 25

ForesightSafety-VLA: A Unified Diagnostic Safety Benchmark for Vision-Language-Action Models

arXiv:2606.2707915.1
Predicted impact top 19% in RO · last 90 daysOriginality Incremental advance
AI Analysis

For researchers developing embodied AI systems, this benchmark provides a systematic way to diagnose safety failures in VLA models, revealing that safety is tightly coupled to perception and control rather than post-hoc filtering.

The paper introduces ForesightSafety-VLA, a diagnostic benchmark for evaluating safety in vision-language-action (VLA) models, defining a 13-category taxonomy and measuring process-level risks. Results show that even the strongest VLA policies incur non-trivial safety costs, with structure and visual variations causing greater safety degradation than language variations.

In embodied intelligence, safety is a prerequisite for reliable robot deployment in the physical world. Current vision-language-action (VLA) models continue to advance toward general-purpose task capability, yet their embodied safety limits remain poorly understood. To address this gap, we introduce ForesightSafety-VLA, a diagnostic benchmark that makes safety the primary evaluation target for VLA systems. We define a 13-category safety taxonomy covering physical interaction safety (Safe-Core), instruction-side safety (Safe-Lang), and perception-side safety (Safe-Vis), and evaluate policies under three controlled dimensions of variation -- scene structure, language command, and visual observation -- so that failure sources can be diagnosed rather than hidden in a single aggregate score. Beyond binary task success, ForesightSafety-VLA measures process-level risk through cumulative safety cost (CC) and risk exposure time (RET), together with a four-quadrant decomposition of safe/unsafe success and failure. We instantiate 66 safety-augmented base scenarios in RoboTwin across 5 embodiments and report results on representative VLA baselines. Across the evaluated baselines, even the strongest policy incurs non-trivial safety cost and unsafe nominal success, while structure and visual variation induce substantially stronger safety degradation than ordinary language variation. These results suggest that embodied safety is tightly coupled to perception, grounding, and control competence rather than being reducible to post-hoc safety filtering alone.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes