CVROJun 17

Motion-Focused Latent Action Enables Cross-Embodiment VLA Training from Human EgoVideos

arXiv:2606.1895515.3
Predicted impact top 25% in CV · last 90 daysOriginality Incremental advance
AI Analysis

This work addresses the problem of leveraging abundant unlabeled human videos for training generalist robot manipulation policies, reducing the need for expensive robot action annotations.

The paper proposes a latent-action framework that extracts general action priors from unlabeled human egocentric videos using a Hybrid Disentangled VQ-VAE, enabling cross-embodiment VLA training. The method achieves competitive performance with state-of-the-art VLA models trained on massive annotated datasets, requiring only 50 trajectories for downstream adaptation.

Training generalist Vision-Language-Action(VLA) models typically requires massive, diverse robotic datasets with high-fidelity action annotations. While egocentric human manipulation videos are abundant and capture significant environmental diversity, the absence of action labels makes them difficult to use in conventional training paradigms. To address this, we propose a latent-action-based framework designed to extract general action priors from unlabeled human videos. The architecture features a Hybrid Disentangled VQ-VAE that decouples motion dynamics from environmental backgrounds through physical masks, enabling the construction of a cross-embodiment action codebook. By pre-training on human videos with the codebook, the VLM backbone learns deep representations of action intent. For adaptation to specific embodiments, we introduce an intent-perception decoupling strategy where the VLM predicts the action intent while a separate frozen visual encoder provides state-specific features to the action expert, thereby reducing action hallucinations. Results in simulation and real-world environments show that our method, pre-trained exclusively on unlabeled human videos, performs competitively with state-of-the-art VLA models trained on massive annotated datasets, requiring only 50 trajectories for downstream adaptation.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes