CVJul 9

ZipDepth: Bringing Lightweight Zero-Shot Monocular Depth Anywhere, on Any Device

arXiv:2607.0877110.5
Predicted impact top 37% in CV · last 90 daysOriginality Incremental advance
AI Analysis

Enables accurate zero-shot depth estimation on resource-constrained devices, addressing the gap between heavy foundation models and domain-limited lightweight methods.

ZipDepth is a compact monocular depth network (6.1M parameters) that achieves real-time performance on devices from servers to embedded platforms, while providing zero-shot generalization competitive with foundation models that are 50x larger, as demonstrated across five benchmarks.

Monocular depth estimation has seen remarkable progress through foundation models achieving robust zero-shot generalization, yet their computational demands place them far beyond the reach of embedded and mobile platforms. Lightweight alternatives exist, but have been developed almost exclusively within single-domain, self-supervised paradigms, failing silently under domain shift. We present ZipDepth, a compact monocular depth network that bridges this gap by combining an efficient reparameterizable encoder-decoder with large-scale knowledge distillation from a foundation model over a large multi-domain training set. Comprising just 6.1M parameters, ZipDepth runs at real-time rates from server GPUs to power-constrained devices, achieving the best trade-off between zero-shot accuracy and deployment efficiency among lightweight models across five benchmarks, taking a significant step towards the accuracy of foundation models with 50x more parameters.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes