ROAIJul 6

Geometry-Aware Motion Latents for Learning Robust Manipulation Policies

arXiv:2607.0471411.7
Predicted impact top 27% in RO · last 90 daysOriginality Highly original
AI Analysis

For robotic manipulation, this provides a more effective action abstraction by focusing on 3D geometric transformations rather than visual appearance, enabling robust policy learning with limited data.

GeoMoLa learns discrete motion latent codes by predicting point cloud evolution during manipulation, achieving state-of-the-art performance with single-view RGB-D input where prior methods need multi-view reconstruction. It succeeds across diverse manipulation benchmarks and enables robust manipulation with minimal demonstrations in cluttered environments.

Learning motion latents for robotic manipulation heavily relies on extracting motion patterns from visual sequences, yet effective action abstractions require understanding three-dimensional geometric transformations. Here, we introduce GeoMoLa (Geometry-Aware Motion Latents), which learns discrete motion latent codes by predicting how point clouds evolve during manipulation rather than reconstructing visual observations. This four-dimensional objective -- spatial geometry changing through time -- forces latent representations to encode actual physical motion rather than appearance patterns. GeoMoLa achieves state-of-the-art performance using only single-view RGB-D input, while existing methods require multi-view reconstruction, succeeding across diverse manipulation benchmarks. Our ablations reveal that geometric prediction is the key to driving performance, quantitatively validating that manipulation depends on spatial understanding. Furthermore, the learned codes exhibit effective motion abstraction: applying them to novel scenes produces physically consistent transformations regardless of visual context. Our real-world experiments also confirm this robustness capability, achieving robust manipulation with minimal demonstrations in cluttered environments where geometric reasoning determines success. Thus, we demonstrate that effective motion latents for robot control can better emerge from understanding motion through its three-dimensional effects rather than pixel-level patterns.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes