CVJun 12

VideoWeave: Unlocking Geometric Consistency in Video Generation via Joint Geometry-Video Modeling

arXiv:2606.14162v17.91 citationsh-index: 3
Predicted impact top 64% in CV · last 90 daysOriginality Incremental advance
AI Analysis

For video generation researchers, this work provides a method to improve geometric consistency without relying on explicit geometry reconstructions, reducing sensitivity to upstream errors.

VideoWeave addresses geometric drift in video generation by jointly modeling geometry and video latents in a shared denoising space, improving geometric coherence while preserving visual quality. Experiments on text-to-video and image-to-video tasks demonstrate enhanced 3D consistency.

Large-scale video diffusion models often fail to preserve 3D structure over time, causing geometric drift and implausible motion under viewpoint changes. Existing methods usually enforce geometric consistency by using explicit geometry reconstructions, such as depth maps, point clouds, or reconstructed 3D structures, to define conditions, supervision, or reward signals, making the generator sensitive to errors from upstream geometry pipelines. We propose VideoWeave, a latent-space post-training framework that uses implicit geometry-model features to constrain the generative distribution, providing a more flexible and non-rigid form of guidance that mitigates the impact of reconstruction errors from geometry models. Specifically, VideoWeave adapts these features into geometry latents and jointly models them with video latents in a shared denoising space, allowing geometry to shape the generative distribution during training. To support this process, we build GeoVid-80K, an 80K-video dataset with paired appearance and geometry representations. Experiments on text-to-video and image-to-video generation show that VideoWeave improves geometric coherence while preserving strong visual quality. VideoWeave project page at https://videoweave.github.io/

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes