ROJun 25

RouterVLA: Turning Smoke Tests into Supervision for Heterogeneous VLA Selection

arXiv:2606.273557.7
Predicted impact top 57% in RO · last 90 daysOriginality Synthesis-oriented
AI Analysis

For robot teams deploying multiple VLA policies, this work provides a practical method to improve system-level performance by reusing rollout data for policy selection, though the gains are incremental and domain-specific.

RouterVLA shows that pre-deployment evaluation rollouts can be reused to supervise policy selection, achieving a +14.64pp gain in held-out success rate (from 0.4686 to 0.6149) across 34,752 LIBERO-Plus records. The study finds that a simple probe-success rule matches learned scorers, and reusing scored trials inflates gains by 1.87x.

We study whether pre-deployment evaluation rollouts can be reused to supervise policy selection. Robot teams routinely smoke test candidate vision-language-action (VLA) policies, then compress those trials into a global winner. RouterVLA evaluates this idea with outcome-disjoint cross-fitting: recorded probes build a profile for each frozen expert, and a separate trial scores the selected expert without entering its profile. Across 34,752 LIBERO-Plus rollout records, a transparent probe-success rule raises held-out success from 0.4686 to 0.6149, a +14.64pp gain. Under the scalar-only profiles studied here, learned scorers are statistically indistinguishable from this rule, showing that commissioning carries the routing value while extra scalar scorer capacity does not create it. Reusing the scored trial inflates the measured gain by $1.87\times$, so credible ledger routing needs outcome separation; model scaling improves individual policies, while commissioning-aware routing improves the system built from them.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes