CVJun 10

4DP-QA: Scalable QA for 4D Perception in Vision Language Models

arXiv:2606.11568v124.11 citationsh-index: 27
Predicted impact top 6% in CV · last 90 daysOriginality Incremental advance
AI Analysis

For researchers working on VLM spatial-temporal reasoning, this work provides a dataset and method to disentangle camera and object motion, addressing a known bottleneck.

VLMs struggle with 4D scene understanding due to indirect motion observation and entangled object/camera motion. The authors propose a QA pipeline with True-Motion Tracking, generating a 400K-sample dataset and 2.2K-sample benchmark, and show that training on their data improves performance on an external benchmark.

Despite recent advances, Vision Language Models (VLMs) still struggle to grasp the dynamics of the world. We note that the ability to reason about a 4D scene, challenging in itself, is further complicated by two factors. First, VLMs observe motion indirectly via its projection onto 2D images. Second, existing datasets fail to disentangle object and camera motion. To address these challenges, we present a QA generation pipeline that focuses on motion-related scene understanding. We take particular care of the entanglement of camera and object motion by casting tracking in both the traditional way and in a novel, fixed reference system, dubbed True-Motion Tracking, which provides an intuitive description of motion. From this pipeline, we generate a large-scale training dataset of 400K samples, 4DP-QA (4D Perception QA), and a 2.2K-sample benchmark, 4DP-QA-Bench. Training existing models on our dataset yields performance improvements on an external benchmark, validating the effectiveness of our method.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes