CVApr 23, 2024

3DBench: A Scalable 3D Benchmark and Instruction-Tuning Dataset

Junjie Zhang, Tianci Hu, Xiaoshui Huang, Yongshun Gong, Dan Zeng

arXiv:2404.14678v15.24 citationsh-index: 16Has CodeIJCAI

Originality Synthesis-oriented

AI Analysis

This provides a more comprehensive evaluation method for researchers in 3D multi-modal AI, though it is incremental as it builds on existing benchmark needs.

The authors tackled the challenge of evaluating Multi-modal Large Language Models (MLLMs) that integrate point cloud and language by introducing 3DBench, a scalable 3D benchmark and instruction-tuning dataset with over 0.23 million QA pairs, which demonstrated superiority in assessing spatial understanding and expressive capabilities.

Evaluating the performance of Multi-modal Large Language Models (MLLMs), integrating both point cloud and language, presents significant challenges. The lack of a comprehensive assessment hampers determining whether these models truly represent advancements, thereby impeding further progress in the field. Current evaluations heavily rely on classification and caption tasks, falling short in providing a thorough assessment of MLLMs. A pressing need exists for a more sophisticated evaluation method capable of thoroughly analyzing the spatial understanding and expressive capabilities of these models. To address these issues, we introduce a scalable 3D benchmark, accompanied by a large-scale instruction-tuning dataset known as 3DBench, providing an extensible platform for a comprehensive evaluation of MLLMs. Specifically, we establish the benchmark that spans a wide range of spatial and semantic scales, from object-level to scene-level, addressing both perception and planning tasks. Furthermore, we present a rigorous pipeline for automatically constructing scalable 3D instruction-tuning datasets, covering 10 diverse multi-modal tasks with more than 0.23 million QA pairs generated in total. Thorough experiments evaluating trending MLLMs, comparisons against existing datasets, and variations of training protocols demonstrate the superiority of 3DBench, offering valuable insights into current limitations and potential research directions.

View on arXiv PDF Code

Similar