Unsupervised Pixel-Level Semantic Left-Right Understanding of In-the-Wild Images
This work addresses the challenging problem of semantic left-right understanding in 2D images for arbitrary objects, which is important for applications like robotics and autonomous driving.
The paper proposes an unsupervised learning framework for pixel-level semantic left-right prediction in single-view images by leveraging 3D shape and image datasets. The method achieves superior performance on both rendered and in-the-wild images compared to existing methods, even for unseen object categories.
While various works address reflective symmetry understanding in 3D data and images, pixel-level semantic left-right prediction of in-the-wild images remains challenging, due to certain difficulties including the lack of 3D information, occlusion, object pose variation, partiality, etc. In this work, we propose an unsupervised learning framework to tackle this challenge. Leveraging recent advances in vertex-wise semantic left-right understanding of 3D data, our unsupervised learning method jointly utilises 3D shape and image datasets to infer pixel-wise semantic left-right predictions in single-view images. In particular, we show that a medium-scale 3D shape dataset comprising mainly of human- and quadruped animal-like shapes, combined with diverse in-the-wild image data, are sufficient to achieve high-quality semantic left-right prediction in images, even for entirely unseen 3D object categories, such as cars or trains. Overall, our approach achieves superior performance in dense pixel-wise semantic left-right predictions on both rendered and in-the-wild image datasets when compared to existing state-of-the-art methods.