Real-World Cooperative Bimanual Dexterous Grasp of Large Objects from Single-View Observations
For robotic manipulation researchers, this work provides a practical solution to real-world bimanual grasping, reducing reliance on full 3D models, though it is an incremental step in a specific domain.
This paper addresses the challenge of cooperative bimanual grasping of large objects in the real world, proposing a framework that generates joint-level grasp configurations from single-view point clouds using a diffusion model, and integrates motion planning with online refinement. Experiments on a dual-arm robot show high success rates on unseen objects, with ablations confirming the contributions of key components.
Bimanual dexterous grasping of large objects is a critical challenge in robotic manipulation. However, most existing studies focus on sequential manipulation rather than cooperative grasping, and methods addressing such bimanual tasks have largely been limited to simulation. These limitations stem from the difficulty of acquiring full 3D object models and generating physically plausible grasping actions. To fill this gap, we propose a real-world bimanual grasping framework that includes: a multimodal dataset capturing joint angles, visual observations and force signals; a Denoising Diffusion Probabilistic Model (DDPM)-based module that generates joint-level grasp configurations from segmented point clouds; and an execution strategy that integrates motion planning with online grasp refinement to ensure physical stability and feasibility. Our approach enables the synthesis of executable bimanual grasps from single-view inputs, reducing dependence on complete 3D object models and ensuring stable real-world performance. Experiments on a dual-arm robot demonstrate high success rates across unseen objects with varying geometries and poses, and ablation studies confirm the contributions of key components of our system.