CVJun 30

InfiniVerse: Occupancy Guided Unbounded Scene Generation for Autonomous Driving

arXiv:2606.3110915.3
Predicted impact top 16% in CV · last 90 daysOriginality Incremental advance
AI Analysis

Provides a unified pipeline for generating realistic, temporally coherent urban environments, addressing a critical need in autonomous driving simulation.

InfiniVerse tackles long-range, controllable urban scene generation for autonomous driving from a single frame, achieving state-of-the-art performance with FID of 6.4 and FVD of 67.97 on Waymo and nuScenes.

Generating realistic, controllable, and temporally coherent urban environments is a critical yet unresolved challenge in the autonomous driving community. In this paper, we introduce InfiniVerse, a unified pipeline for long-range, 2D-3D-aligned, and controllable synthesis of dynamic urban scenes from a single frame. In practice, our approach first reconstructs a 3D occupancy representation from the input multi-view frame. This representation serves as a foundation for autoregressive scene extension along arbitrary trajectories. Subsequently, a video diffusion model translates the coarse occupancy grid into realistic, spatiotemporally consistent video sequences. Moreover, we propose a hierarchical sketch-and-refine paradigm, in which the generated videos are re-projected as image-conditioned feedback to enhance the 3D occupancy representation, establishing cross-modal alignment and mutual enhancement between the visual and spatial domains. Extensive evaluations on the Waymo Open Dataset and nuScenes demonstrate that InfiniVerse achieves state-of-the-art performance, with a FID of 6.4 and FVD of 67.97, significantly outperforming existing benchmarks in both duration and stability.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes