CVAIGRJul 1

World from Motion: Generative Dynamic Gaussian Reconstruction from Monocular Video

arXiv:2607.0120217.0
Predicted impact top 13% in CV · last 90 daysOriginality Highly original
AI Analysis

This work addresses the problem of reconstructing dynamic 3D scenes from monocular video, a key challenge in computer vision and graphics, with significant improvements in novel-view synthesis and 3D motion accuracy.

World from Motion generates dynamic 3D Gaussian representations from monocular videos, achieving state-of-the-art 4D reconstruction and generalizing to in-the-wild videos with large viewpoint changes and dynamic motions.

We present World from Motion, a method for generating freely renderable dynamic 3D Gaussian representations from monocular videos. Our approach conditions a video model on dense, pixel-aligned renderings that encode appearance, geometry, and 3D scene motion along both input and target camera trajectories to correct rendering artifacts and fill in missing regions from an initial reconstruction. To train this model, we construct a dataset of aligned multiview video pairs and dynamic 3DGS representations, with simulated artifacts characteristic of monocular reconstruction. At test time, we distill the model's generations, including newly observed regions and motions, back into a single consistent, high-quality dynamic 3DGS, improving both novel-view synthesis and the underlying 3D motion. Our method sets a new state of the art in 4D reconstruction and seamlessly generalizes to in-the-wild videos with large viewpoint changes and dynamic motions.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes