VLRC: Vision-Language Reprojection Consistency as a scalable signal for better feed-forward 3D pretraining
For researchers in 3D vision and scene understanding, VLRC provides a scalable way to leverage vision-language models to improve 3D pretraining without additional annotations, addressing the incomplete learning signals of existing methods.
VLRC introduces a scalable auxiliary objective that uses frozen vision-language representations as semantic multi-view supervision for feed-forward 3D pretraining, improving depth and camera estimation as well as enabling coherent multi-view semantic fusion. Experiments show consistent gains in 3D reconstruction accuracy and zero-shot open-vocabulary 3D semantic segmentation on indoor and outdoor benchmarks.
Feed-forward 3D models are commonly trained using either expensive geometric supervision or self-supervised photometric objectives, both of which provide incomplete learning signals. We introduce Vision-Language Reprojection Consistency (VLRC), a scalable auxiliary objective that exploits frozen vision-language representations as semantic multi-view supervision. Given a predicted 3D reconstruction, VLRC reprojects dense vision-language features across views and enforces feature consistency between corresponding image locations, requiring no additional 3D annotations. The objective integrates seamlessly with both self-supervised monocular reconstruction and supervised-pretrained feed-forward 3D models during unlabeled adaptation. By aligning geometry with language-grounded features, VLRC not only improves depth and camera estimation but also enables more coherent multi-view semantic fusion for open-vocabulary 3D scene understanding. Experiments on indoor and outdoor benchmarks demonstrate consistent gains in 3D reconstruction accuracy and zero-shot open-vocabulary 3D semantic segmentation.