CV CLJan 12, 2025

3DCoMPaT200: Language-Grounded Compositional Understanding of Parts and Materials of 3D Shapes

Mahmoud Ahmed, Xiang Li, Arpit Prajapati, Mohamed Elhoseiny

arXiv:2501.06785v14 citationsh-index: 24Has Code

Originality Synthesis-oriented

AI Analysis

This work addresses the problem of fine-grained compositional understanding of 3D shapes for researchers and applications in robotics and AI, but it is incremental as it builds upon and expands an existing dataset.

The authors tackled the limited scope of existing datasets for part-level 3D object understanding by introducing 3DCoMPaT200, a large-scale dataset with 200 object categories, 1,031 part categories, and 293 material classes, which is about 5 times larger in object vocabulary and 4 times larger in part categories compared to prior datasets, and they demonstrated its effectiveness through a compositional part shape retrieval task where model performance improved with more parts described in text.

Understanding objects in 3D at the part level is essential for humans and robots to navigate and interact with the environment. Current datasets for part-level 3D object understanding encompass a limited range of categories. For instance, the ShapeNet-Part and PartNet datasets only include 16, and 24 object categories respectively. The 3DCoMPaT dataset, specifically designed for compositional understanding of parts and materials, contains only 42 object categories. To foster richer and fine-grained part-level 3D understanding, we introduce 3DCoMPaT200, a large-scale dataset tailored for compositional understanding of object parts and materials, with 200 object categories with $\approx$5 times larger object vocabulary compared to 3DCoMPaT and $\approx$ 4 times larger part categories. Concretely, 3DCoMPaT200 significantly expands upon 3DCoMPaT, featuring 1,031 fine-grained part categories and 293 distinct material classes for compositional application to 3D object parts. Additionally, to address the complexities of compositional 3D modeling, we propose a novel task of Compositional Part Shape Retrieval using ULIP to provide a strong 3D foundational model for 3D Compositional Understanding. This method evaluates the model shape retrieval performance given one, three, or six parts described in text format. These results show that the model's performance improves with an increasing number of style compositions, highlighting the critical role of the compositional dataset. Such results underscore the dataset's effectiveness in enhancing models' capability to understand complex 3D shapes from a compositional perspective. Code and Data can be found at http://github.com/3DCoMPaT200/3DCoMPaT200

View on arXiv PDF Code

Similar