SeeSE3: Emergence of 3D Space in Vision Features
For researchers in computer vision and representation learning, this work reveals that self-supervised vision models inherently encode 3D spatial structure, enabling new latent-space navigation techniques.
The paper investigates whether vision foundation models encode 3D Euclidean space properties in their latent representations, finding that self-supervised models possess latent subspaces strongly correlated with 3D space. This insight enables latent-space visual odometry and localization without explicit 3D reconstruction.
In this paper, we ask whether vision foundation models construct representations that reflect the intrinsic properties of 3D Euclidean space. Unlike previous works that probe 3D awareness of vision features by regressing image-centric quantities such as depth or normals, we investigate the relation between the structure of the space of visual features and the group of Euclidean transformations $SE(3)$. We propose a set of probes to evaluate this relation from both topological and geometric perspectives: a mutual neighborhood metric that measures the alignment between feature neighborhoods and spatial topology, and a Poincaré Adapter to test the linear accessibility of the geometry of camera motion from latent displacements in static scenes. We show that self-supervised vision models, which, in principle, have not been trained with direct 3D supervision or active agency, possess latent subspaces that are remarkably strongly correlated with three-dimensional Euclidean space, when probed correctly. Building on this insight we propose a new class of "Latent-Space Navigation" techniques that perform visual odometry and localization purely in the latent space, bypassing the need for explicit 3D reconstruction.