CVAIJul 9

MultiView-Bench: A Diagnostic Benchmark for World-Centric Multi-View Integration in VLMs

arXiv:2607.089708.3h-index: 3
Predicted impact top 52% in CV · last 90 daysOriginality Incremental advance
AI Analysis

Provides a much-needed diagnostic tool for a core cognitive ability (multi-view integration) in VLMs, revealing consistent failure modes that limit their deployment in 3D tasks.

MultiView-Bench is a diagnostic benchmark for evaluating VLMs' ability to integrate multiple viewpoints into a coherent 3D world model. Frontier VLMs show strong 2D performance but struggle with 3D spatial relations and multi-view aggregation; the proposed ViewNavigator framework improves performance by 3-5x.

Recent benchmarks for VLMs largely assess single- or limited-view perception, leaving untested the core cognitive ability to integrate observations across viewpoints into a coherent, world-centric (allocentric) 3D mental model. We introduce MultiView-Bench, a diagnostic benchmark expressly designed to evaluate multi-view integration for holistic 3D scene comprehension. Unlike existing datasets that focus on pixel-level mapping or camera-relative navigation, MultiView-Bench requires models to decouple object positioning from transient perspectives and ground them in a fixed global coordinate system. This capability serves as a prerequisite for VLMs before being deployed for downstream tasks such as mechanical part assembly. Our systematic evaluation of frontier VLMs reveals consistent failure modes: strong performance on 2D planar relations from a single image, but marked difficulty with 3D spatial relations and with aggregating information across views. We further identify biases in VLMs, such as struggles with unconventional axis directions and sensitivity to object colorways and texture variations. Acknowledging these limitations, we propose ViewNavigator, a multi-agent framework that actively selects informative viewpoints, perceives, and fuses multi-view evidence, improving diverse base models on MultiView-Bench even under a strict budget-matched comparison (and by 3-5x for the full agent).

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes