ROCVAug 6

XEWorld: Can Action-Conditioned World Models Generalize to Unseen Robot Embodiments?

arXiv:2608.0579921.2h-index: 7
Predicted impact top 4% in RO · last 90 daysOriginality Incremental advance
AI Analysis

For robotic manipulation researchers, this work identifies a critical limitation in world models, showing that they lack true physical understanding, which is essential for generalizing to new robots.

The paper introduces XEWorld, a cross-embodiment testbed for action-conditioned world models, and finds that current models generalize based on visual similarity rather than physical kinematics, failing to render unseen embodiments zero-shot and suffering catastrophic forgetting in few-shot adaptation.

Action-conditioned world models are promising learned simulators for robotic manipulation, yet evaluating them exclusively on training robots fails to reveal whether they capture physical dynamics or merely memorize visual patterns. To answer whether a model can faithfully render a robot it has never seen, we introduce XEWorld, a controlled cross-embodiment testbed for world models that isolates embodiments by evaluating held-out robots within physically identical scenes. Our systematic analysis uncovers a shared architectural bottleneck: current models act primarily as 2D visual pattern matchers whose generalization is governed by visual similarity rather than physical kinematic similarity. Driven by this limitation, they struggle to translate abstract numeric joint actions into coherent visual trajectories, and fail to predict dynamic visual changes from static initial observations. Consequently, successfully rendering an unseen embodiment zero-shot strictly requires heavily grounded cues, specifically pixel-space actions and explicit spatial-temporal alignment. Even when bypassing this zero-shot barrier via few-shot adaptation, the forced appearance recovery triggers catastrophic forgetting of seen embodiments. Together, these failures expose a critical inability to apply learned physical dynamics to novel visual appearances, highlighting that achieving true cross-embodiment generalization requires architectural innovations that decouple visual appearance from underlying physical dynamics.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes