CVJul 21

IGGT4D: Streaming 4D Instance-Grounded Geometry Transformer

arXiv:2607.1922818.0
Predicted impact top 8% in CV · last 90 daysOriginality Incremental advance
AI Analysis

This work addresses the lack of temporally consistent object-level understanding in streaming 3D reconstruction, which is crucial for real-world spatial intelligence applications like robotics and autonomous driving.

IGGT4D is a streaming Transformer for online 4D scene understanding that jointly reconstructs geometry and instance identities from video streams, achieving state-of-the-art performance on 3D reconstruction, pose estimation, instance tracking, and open-vocabulary segmentation while maintaining scalable online inference for long dynamic sequences.

Real-world spatial intelligence requires agents to understand scenes from continuous video streams, where objects move, persist, disappear, and reappear over time. While recent spatial foundation models have enabled generalizable feed-forward 3D reconstruction, most streaming methods remain geometry-centric and lack temporally consistent object-level understanding. Meanwhile, existing semantic reconstruction and 3D-aware vision-language methods largely rely on externally extracted 2D semantic cues or loosely coupled geometry inputs, limiting unified geometry-instance learning in long dynamic scenes. In this paper, we propose IGGT4D, a streaming instance-grounded geometry Transformer for online 4D scene understanding. IGGT4D processes video frames sequentially, reuses historical context through causal spatial-temporal modeling, and incrementally updates a unified representation of camera motion, geometry, and object identity. This enables long-sequence feed-forward reconstruction with geometry-instance consistency in dynamic environments. To address the lack of high-quality 4D supervision, we further construct InsScene4D-147K, a large-scale dataset spanning real/synthetic and static/dynamic scenes, with RGB images, depth, poses, and temporally consistent instance masks generated by an automated geometry-guided annotation pipeline. Experiments on 3D reconstruction, pose estimation, instance spatial tracking, and open-vocabulary segmentation demonstrate that IGGT4D outperforms existing streaming baselines while maintaining scalable online inference for long dynamic sequences.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes