xperception -- Making Robotic Grasping Easier
For high-mix low-volume manufacturing, xperception removes the bottleneck of retraining vision systems for new objects, enabling flexible robotic grasping.
xperception introduces a zero-shot 6D pose estimation technology that eliminates object-specific fine-tuning and data annotation, achieving millimeter-accurate pose estimation using CAD models and foundation models. It won the BOP Challenge 2024 and is validated at TRL 6 for industrial bin picking.
The transition toward high-mix low-volume manufacturing demands flexibility in robotic manipulation. However, conventional vision systems remain a bottleneck, requiring extensive data collection and model retraining whenever a new object is introduced to the production line. To overcome this rigidity, we present xperception, a zero-shot 6D pose estimation technology that eliminates the need for object-specific fine-tuning and laborious data annotation. By directly utilizing typical CAD models and integrating the rich semantic features of foundation models (e.g. DINOv2, GeDi), xperception achieves millimeter-accurate 6D pose estimation. xperception showed robustness against severe occlusions in industrial tasks like bin picking and is engineered for deployment on industrial edge hardware, such as NVIDIA Jetson Thor. Validated at a TRL of 6, the core methodology behind xperception is based on the FreeZe algorithm, which won the international BOP Challenge 2024, paving the way for scalable, plug-and-play robotic automation in unstructured high-mix low-volume manufacturing industries.