Scaling Cross-Embodiment World Models for Dexterous Manipulation
This work addresses the challenge of data sharing and control transfer across diverse robot and human hand morphologies, offering a scalable approach for generalist robot learning.
The paper introduces a particle-based world model that abstracts embodiment-specific joint spaces into a shared geometric representation, enabling cross-embodiment learning for dexterous manipulation. Experiments show that increasing embodiment diversity improves generalization, combining simulated and real data outperforms either alone, and the model enables effective control on robotic hands with different kinematics.
Cross-embodiment learning seeks to build generalist robots that learn from and operate across diverse morphologies, but differences in kinematics and action spaces hinder data sharing and control transfer. We ask: What structure can be shared across embodiments despite these differences? We argue that the physical interactions they induce can be modeled in a shared geometric space, allowing world models to provide a common interface for learning and control. To realize this idea, we represent human and robot hands as sets of 3D particles and define actions as end-effector particle displacement fields. This representation abstracts away embodiment-specific joint spaces while preserving the geometry and motion relevant to physical interaction. We train a graph-based world model on random interaction data from diverse simulated robot hands and real human hands, and integrate it with model-predictive control for deployment on new hardware. Experiments on rigid and deformable manipulation reveal three findings: increasing the diversity of training embodiments improves generalization to unseen hands; appropriately combining simulated and real-world data outperforms either source alone; and the same learned model enables effective control on robotic hands with distinct kinematics and degrees of freedom. These results position particle-based world models as a shared interface for learning from and for heterogeneous embodiments.