CVJul 6

Vision Pretraining for Dense Spatial Perception

arXiv:2607.0524710.6
Predicted impact top 36% in CV · last 90 daysOriginality Incremental advance
AI Analysis

This work addresses the need for dense spatial perception in embodied AI by introducing a scalable pretraining principle that enhances geometric understanding, though it is incremental over existing self-supervised methods like DINOv3.

The authors propose masked boundary modeling, a self-supervised pretraining paradigm that learns sub-pixel boundary representations and uses them as masked targets for dense visual token learning. Their method, LingBot-Vision, improves depth estimation in LingBot-Depth 2.0 over the previous version, demonstrating efficacy across diverse downstream tasks.

Dense spatial perception is essential for physical intelligence, where visual systems are expected to recover structured, metric, and actionable representations from pixel observations. Modern visual foundation models tend to prioritize semantic invariance, often at the expense of detailed spatial understanding. In this work, we study vision pretraining through a boundary-centric lens, motivated by the premise that boundaries and shape discontinuities offer essential cues for perceiving geometric properties. Concretely, we propose masked boundary modeling, a self-supervised paradigm that dynamically learns sub-pixel boundary representations and subsequently leverages the discovered boundary-bearing tokens as masked targets to facilitate dense visual token learning. By scaling this framework, we develop LingBot-Vision and demonstrate its efficacy across a diverse set of downstream vision tasks with DINOv3 as a strong baseline. Remarkably, LingBot-Vision drives the progression from LingBot-Depth 1.0 to LingBot-Depth 2.0 for depth completion, and thereby yields enhanced depth estimation, a key pillar for embodied artificial intelligence. Our findings reveal that boundary modeling goes beyond simple line segments and instead serves as a scalable pretraining principle for learning spatially structured visual representations.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes