ROCVJun 15

V2P-Manip: Learning Dexterous Manipulation from Monocular Human Videos

arXiv:2606.164369.9
Predicted impact top 41% in RO · last 90 daysOriginality Incremental advance
AI Analysis

For robotics researchers, this provides a scalable alternative to costly teleoperation data for learning dexterous manipulation, though the method is incremental.

V2P-Manip learns dexterous manipulation policies from monocular human videos, achieving over 75% average success rate across multiple synthetic tasks and outperforming prior methods on TACO and OakInk benchmarks.

Achieving autonomous robotic dexterous manipulation requires precise, human-like action sequences at scale. As a scalable supplement to costly teleoperation data, extracting trajectories with both visual fidelity and physical plausibility from monocular videos represents a promising frontier in embodied AI. To this end, we introduce V2P-Manip, an efficient framework designed to learn dexterous manipulation policies directly from human demonstration videos. We establish an efficient, integrated pipeline encompassing 3D asset acquisition, trajectory estimation, and dexterous policy learning. To bridge the gap between visual perception and physical constraints, we introduce a two-stage refinement process to enforce spatial alignment and physical consistency. Evaluations on the TACO and OakInk benchmarks demonstrate that our approach significantly outperforms previous methods in pose accuracy, adaptability to unstructured environments, and training efficiency. Ultimately, experimental results confirm an average success rate of over 75% across multiple synthetic manipulation tasks and validate the adaptability of the extracted manipulation priors across diverse dexterous hand embodiments.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes