CVJun 12

MooMIns -- Monocular 3D Reconstruction and Object Pose Estimation from Multiple Instances

arXiv:2606.14389v17.7
Predicted impact top 65% in CV · last 90 daysOriginality Incremental advance
AI Analysis

For industrial bin-picking applications, this method enables geometry-based 3D reconstruction and pose estimation from a single image without relying on learned depth priors.

MooMIns exploits multiple instances of an object in a single monocular image to perform simultaneous 3D reconstruction and 6D pose estimation, achieving accurate reconstruction of unseen objects and reliable pose estimation in bin-picking scenarios.

Simultaneous 3D reconstruction and 6D object pose estimation from a single monocular image is an inherently ill-posed problem. In industrial settings, however, multiple instances of an object are often randomly arranged in bins, implicitly providing several views of the same object within a single image. We show that this implicit multi-view geometry can be exploited to simultaneously reconstruct the object in 3D and estimate the 6D pose of each visible object instance. We present MooMIns, a new Gaussian-splatting-based approach that inverts the original Gaussian splatting formulation: instead of rendering a single scene from multiple cameras, we render multiple object instances from a single camera. Our method is initialized with SAM3 instance segmentation masks and a modified Structure from Motion (SfM) pipeline. In contrast to learned monocular depth estimation, we perform true geometry-based reconstruction from image evidence, avoiding hallucinations caused by training data priors. We evaluate MooMIns on synthetic and real bin-picking scenarios, and demonstrate accurate reconstruction of previously unseen objects as well as reliable pose estimation of individual instance

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes