ROJun 18

Perturbation-Based Uncertainty for Failure Detection in Vision-Language-Action Models

arXiv:2606.2075411.6
Predicted impact top 31% in RO · last 90 daysOriginality Incremental advance
AI Analysis

It addresses the need for reliable uncertainty quantification in robotic manipulation VLA models, which lack explicit predictive probabilities, offering a practical solution for failure detection.

The paper proposes a label-free, model-agnostic uncertainty estimation method for VLA models by perturbing hidden activations, improving failure detection under distribution shift over sampling-based methods on LIBERO and LIBERO-PRO benchmarks.

Vision-Language-Action (VLA) models have shown strong performance in robotic manipulation, but reliable uncertainty quantification remains challenging, particularly under distribution shift. Unlike autoregressive policies, many modern VLA models generate continuous actions through regression or flow-based generation, where explicit predictive probabilities are unavailable. Moreover, existing approaches often rely on stochastic action sampling or supervised failure labels, limiting their applicability across diverse pretrained VLA models. In this work, we propose a label-free and model-agnostic framework for inference-time uncertainty estimation through hidden activation perturbations, motivated by Bayesian perspectives on local model variations. Specifically, we inject Gaussian perturbations into transformer hidden activations and estimate epistemic signals from disagreement across perturbed action predictions. Experiments on LIBERO and LIBERO-PRO show that perturbation-based uncertainty consistently improves failure detection under distribution shift compared to sampling-based uncertainty, providing a practical uncertainty signal for VLA models.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes