ROJul 13

Replanning Human-Robot Collaborative Tasks with Vision-Language Models via Semantic and Physical Dual-Correction

arXiv:2602.145518.41 citationsh-index: 26
Predicted impact top 42% in RO · last 90 daysOriginality Incremental advance
AI Analysis

For human-robot collaboration researchers, this work addresses the challenge of integrating VLM-based semantic reasoning with physically executable motion planning, but the results are incremental and limited by visual verification issues.

The paper proposes a replanning framework for human-robot collaborative assembly that uses vision-language models with dual-correction mechanisms (internal logical verification and external visual verification) to handle ambiguous corrective instructions. Real-world experiments achieved 66.7% success in object fixation, 100% in initial tool selection, and 75.0% in corrective tool selection, though visual verification was identified as a key limitation.

Human-robot collaborative assembly requires robots to interpret ambiguous corrective instructions while producing physically executable motions. Vision-language models (VLMs) provide semantic reasoning but may select logically inconsistent targets or misjudge execution outcomes. We propose a replanning framework that maps human instructions to Action Target candidates, including grasp poses and tool selections, and combines an Internal Correction Model for pre-execution logical verification with an External Correction Model for post-execution visual verification. The framework integrates VLM reasoning with 6-DoF grasp generation and collision-free trajectory planning. Simulation ablations show configuration-dependent effects: internal correction improves candidate validity, whereas external correction enables recovery for a low-latency VLM but can reduce success when visual verification produces false negatives. Experiments with an upper-body humanoid robot achieved 66.7% success in real-world object fixation, 100% in initial tool selection, and 75.0% in corrective tool selection. These results demonstrate interactive replanning across spatial and semantic collaborative tasks while identifying visual-state verification as a key limitation.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes