CVJun 14

3D Consistency Optimization for Self-Supervised Monocular Video Depth Estimation

arXiv:2606.1568111.6
Predicted impact top 40% in CV · last 90 daysOriginality Highly original
AI Analysis

It addresses the problem of geometrically inconsistent depth predictions and cross-frame drift in self-supervised video depth estimation, which is crucial for applications like endoscopic navigation.

This paper introduces a new paradigm for self-supervised monocular video depth estimation that reformulates the problem as multi-view 3D reconstruction, achieving state-of-the-art spatial accuracy and outperforming existing frame-based, video-based, and multi-view baselines.

Reliable monocular video depth estimation is crucial for downstream 3D reasoning and embodied AI in endoscopic navigation. However, existing self-supervised approaches typically treat video frames independently or rely on weak temporal regularization. These methods, lacking a holistic perception of the underlying 3D scene, inevitably suffer from geometrically inconsistent predictions and severe cross-frame drift. To address these limitations, we introduce a new paradigm that recasts sequential video depth estimation as an unconstrained multi-view 3D reconstruction problem, enabling full exploitation of the powerful geometric priors embedded in recent 3D foundation models. The core of our approach is a 3D consistency optimization framework driven by three constraints: image-level photometric rendering, explicit world-coordinate geometric alignment, and multi-scale temporal gradient consistency. Such unified optimization elegantly anchors isolated frames to a globally coherent 3D structure. Our method has been validated in both the self-supervised training scenarios and challenging zero-shot clinical environments. Results show that the proposed approach achieves state-of-the-art spatial accuracy, outperforming the frame-based, video-based depth estimators and the multi-view 3D reconstruction baselines.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes