CVMar 10, 2025

Aligning Instance-Semantic Sparse Representation towards Unsupervised Object Segmentation and Shape Abstraction with Repeatable Primitives

Jiaxin Li, Hongxing Wang, Jiawei Tan, Zhilong Ou, Junsong Yuan

arXiv:2503.06947v18.42 citationsh-index: 3IEEE Trans Vis Comput Graph

Originality Incremental advance

AI Analysis

This work addresses the need for efficient shape reasoning in computer vision by reducing reliance on supervised data and multi-stage training, though it appears incremental in improving unsupervised methods.

The paper tackles the problem of unsupervised 3D object shape understanding by introducing a one-stage framework that jointly performs instance segmentation, semantic segmentation, and shape abstraction, achieving this through sparse representation and feature alignment without costly annotations.

Understanding 3D object shapes necessitates shape representation by object parts abstracted from results of instance and semantic segmentation. Promising shape representations enable computers to interpret a shape with meaningful parts and identify their repeatability. However, supervised shape representations depend on costly annotation efforts, while current unsupervised methods work under strong semantic priors and involve multi-stage training, thereby limiting their generalization and deployment in shape reasoning and understanding. Driven by the tendency of high-dimensional semantically similar features to lie in or near low-dimensional subspaces, we introduce a one-stage, fully unsupervised framework towards semantic-aware shape representation. This framework produces joint instance segmentation, semantic segmentation, and shape abstraction through sparse representation and feature alignment of object parts in a high-dimensional space. For sparse representation, we devise a sparse latent membership pursuit method that models each object part feature as a sparse convex combination of point features at either the semantic or instance level, promoting part features in the same subspace to exhibit similar semantics. For feature alignment, we customize an attention-based strategy in the feature space to align instance- and semantic-level object part features and reconstruct the input shape using both of them, ensuring geometric reusability and semantic consistency of object parts. To firm up semantic disambiguation, we construct cascade unfrozen learning on geometric parameters of object parts.

View on arXiv PDF

Similar