ROLGJun 16

Uncertainty Quantification for Flow-Based Vision-Language-Action Models

arXiv:2606.1804315.6
Predicted impact top 18% in RO · last 90 daysOriginality Incremental advance
AI Analysis

For roboticists deploying VLAs in non-stationary environments, this provides a practical method to detect unreliable predictions and reduce costly data collection.

This work addresses the lack of uncertainty quantification in flow-based vision-language-action models (VLAs) for robotic manipulation. By proposing a velocity-field disagreement (VFD) method and the SAVE framework for active fine-tuning, they achieve at least 22% fewer expert demonstrations needed for adaptation and demonstrate improved failure detection on the LIBERO benchmark.

Vision-language-action models (VLAs) combine vision-language backbones with expressive generative action heads trained via flow matching on large-scale robotic datasets. Despite their strong empirical performance in robotic manipulation, VLAs lack mechanisms to quantify confidence in their predictions and to detect when their actions may be unreliable. This presents a critical limitation for real-world deployment in non-stationary environments, where models inevitably encounter scenarios outside their pretraining distribution and may fail without warning. To address this, we derive an efficient method for quantifying epistemic uncertainty in flow-matching models by leveraging velocity-field disagreement (VFD) across a small ensemble. We successfully use this uncertainty estimate for failure detection during deployment and active fine-tuning of flow-based VLAs. To this end, we propose SAVE, a framework for uncertainty-guided active multitask fine-tuning that reduces the number of costly expert demonstrations required to adapt VLAs to new tasks. Through extensive experiments on the LIBERO benchmark, we demonstrate that VFD yields better-calibrated uncertainty estimates predictive of downstream performance, that VFD achieves strong performance in detecting failures, and that uncertainty-guided data acquisition with SAVE requires at least 22% fewer samples than baselines. In summary, our work shows that quantifying epistemic uncertainty in flow-based VLAs improves both failure awareness and adaptation. Project website: tum-lsy.github.io/uq_vla/.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes