ROLGJun 23

RE4: Transformation-aware Imitation of Object Interactions Using Manipulation Modes

arXiv:2606.244035.9
Predicted impact top 68% in RO · last 90 daysOriginality Incremental advance
AI Analysis

For robotic imitation learning, RE4 provides an interpretable alternative to end-to-end methods that maintains high performance, particularly in low-data and adversarial settings.

RE4 introduces a framework for imitation learning of object interactions that combines self-supervised pose estimation, mode-aware retrieval and transformation, and replanning, achieving robust performance in sparse data regimes. On image-based Push-T, RE4 achieves 0.95 success rate with 10 demonstrations, outperforming diffusion-based methods by 15%.

Object interaction tasks have been a focus of advances in imitation learning. End-to-end methods, dominated by diffusion and flow-based variants have shown leaps in performance while sacrificing interpretability. Object-centric and pose-informed variants have had a role in learning from demonstration in manipulation tasks. In this paper, we revisit a few modern imitation learning benchmarks for object interactions, with the aim of composing a framework that repurposes principled theories of manipulation, preserving both performance and interpretability. For image observations, lightweight training is proposed for model-free pose estimation of the target object, using self-supervision over the demonstration data available for imitation learning. This information is then used to inform a manipulation mode-aware retrieval of a demonstration, a mode-aware transformation, a replan step that connects to the retrieval point while preserving mode constraints, and finally rolling out the transformed demonstration. These compose four key steps of the proposed RE4 framework, evaluated over state-based and image-based benchmarks in Push-T and Robomimic. An adversarial benchmark that evaluates sparse data regions of image-based Push-T showcases the robustness, further bolstered by indications from low-data regime experiments. The current work shows promise in using simple interpretable building blocks to learn manipulation skills.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes