CVJul 2

VLRC: Vision-Language Reprojection Consistency as a scalable signal for better feed-forward 3D pretraining

arXiv:2607.027079.3
Predicted impact top 45% in CV · last 90 daysOriginality Incremental advance
AI Analysis

For researchers in 3D vision and scene understanding, VLRC provides a scalable way to leverage vision-language models to improve 3D pretraining without additional annotations, addressing the incomplete learning signals of existing methods.

VLRC introduces a scalable auxiliary objective that uses frozen vision-language representations as semantic multi-view supervision for feed-forward 3D pretraining, improving depth and camera estimation as well as enabling coherent multi-view semantic fusion. Experiments show consistent gains in 3D reconstruction accuracy and zero-shot open-vocabulary 3D semantic segmentation on indoor and outdoor benchmarks.

Feed-forward 3D models are commonly trained using either expensive geometric supervision or self-supervised photometric objectives, both of which provide incomplete learning signals. We introduce Vision-Language Reprojection Consistency (VLRC), a scalable auxiliary objective that exploits frozen vision-language representations as semantic multi-view supervision. Given a predicted 3D reconstruction, VLRC reprojects dense vision-language features across views and enforces feature consistency between corresponding image locations, requiring no additional 3D annotations. The objective integrates seamlessly with both self-supervised monocular reconstruction and supervised-pretrained feed-forward 3D models during unlabeled adaptation. By aligning geometry with language-grounded features, VLRC not only improves depth and camera estimation but also enables more coherent multi-view semantic fusion for open-vocabulary 3D scene understanding. Experiments on indoor and outdoor benchmarks demonstrate consistent gains in 3D reconstruction accuracy and zero-shot open-vocabulary 3D semantic segmentation.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes