CVROJul 1

Privacy-Preserving Depth-Only Open-Vocabulary 3D Semantic Segmentation Via Uncertainty-Guided Test-Time Optimization

arXiv:2607.0097811.9
Predicted impact top 30% in CV · last 90 daysOriginality Incremental advance
AI Analysis

For 3D scene understanding in privacy-sensitive indoor environments, UTTO enables open-vocabulary segmentation without RGB data, reducing privacy risks while maintaining performance.

UTTO addresses privacy-preserving open-vocabulary 3D semantic segmentation using only depth data, converting prediction uncertainty into guidance for test-time optimization. It consistently improves depth-only segmentation on ScanNet20, ScanNet40, and ScanNet200, outperforming baselines under privacy constraints.

Privacy-preserving perception is a critical requirement for deploying 3D scene understanding systems in real-world indoor environments, yet it remains underexplored in open-vocabulary 3D semantic segmentation. Existing methods typically rely on obtaining rich semantic cues from RGB images, which may expose privacy-sensitive visual information. Depth-only 3D geometry provides a privacy-preserving alternative, but the absence of appearance-based semantic cues makes open-vocabulary predictions highly uncertain and less reliable. Under this setting, we propose to convert uncertainty into a guidance signal to identify unreliable semantic responses and use semantic priors from foundation models to regularize their refinement. We present UTTO, an uncertainty-guided test-time optimization framework for depth-only open-vocabulary 3D semantic segmentation. Without additional training, experiments on ScanNet20, ScanNet40, and ScanNet200 demonstrate that UTTO consistently improves depth-only open-vocabulary 3D segmentation and outperforms representative baselines under privacy-preserving conditions.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes