ROJun 22

Cloak: Zero-Shot Cross-Embodiment Manipulation by Masking the End-Effector from the VLA

arXiv:2606.2283616.0
Predicted impact top 17% in RO · last 90 daysOriginality Incremental advance
AI Analysis

This work addresses the problem of robotic manipulation across different hardware embodiments, offering a practical solution for data reuse and generalization.

Cloak enables zero-shot cross-embodiment transfer for VLA models by masking the end-effector from wrist camera views, achieving successful transfer to unseen grippers, arms, and a five-fingered hand without any new data collection.

We present Cloak, a training recipe that endows a Vision-Language-Action (VLA) model with zero-shot cross-embodiment transfer by cloaking the end-effector from its own wrist camera. The end-effector occupies a large and consistent region of the wrist view and masking it allows for embodiment-agnostic visual reasoning. Cloak renders a mask in simulation from the robot's known geometry, accurately and in real time, with no segmentation or generative models. During training, we augment the mask so the model generalizes to embodiments unseen at training time. We demonstrate the recipe with Cloak-VLA, a VLA trained with Cloak on a single parallel-jaw gripper dataset. No data of new embodiments is ever collected. Cloak-VLA transfers zero-shot to various unseen embodiments, including another gripper, another arm, and a five-fingered hand, while preserving the source embodiment's performance. By decoupling the wrist view from its own embodiment, Cloak allows data to outlive the hardware it was collected on.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes