GWM-VLA: Geometry-Aware Latent World Modeling for Vision-Language-Action Learning
This work aims to improve the robustness of robotic manipulation for VLA models, which is an incremental improvement for researchers and practitioners in robotics.
This paper addresses the degradation of Vision-Language-Action (VLA) models in robotic manipulation under visual and environmental shifts by proposing GWM-VLA, a geometry-aware latent world modeling framework. GWM-VLA uses geometry-aware multi-view state encoding and global context-conditioned target-view prediction, demonstrating effectiveness and robustness in both simulation and real-world environments.
Vision-Language-Action (VLA) models achieve strong robotic manipulation performance but often degrade under visual and environmental shifts. Latent world modeling offers a promising approach to improving robustness, yet existing methods commonly encode camera views independently and predict holistic scene dynamics without explicitly modeling their geometric relationships. We propose GWM-VLA, a geometry-aware latent world modeling framework for VLA learning. GWM-VLA combines geometry-aware multi-view state encoding, global context-conditioned target-view prediction, and shared latent-action representations grounded by robot-action supervision. Specifically, VGGT-$Ω$ jointly aggregates multi-view observations at each timestep to construct geometry-aware multi-view states. The latent world model predicts the next-step patch tokens of a selected target view using patch and register tokens obtained after multi-view aggregation, thereby retaining multi-view geometric information without predicting the complete multi-view state. We use the wrist view as the target in our experiments, placing greater emphasis on end-effector motion and local gripper-object interactions. Finally, the shared latent-action representations condition both the latent world model and the flow-matching action head, allowing latent-prediction supervision and ground-truth robot-action supervision to jointly shape the same latent-action representations. Experiments across both simulation and real-world environments demonstrate the effectiveness and robustness of GWM-VLA.