CasaMaestro: Multi-View Panoramas for House-Scale 3D Reconstruction
This work addresses the need for efficient, metric 3D reconstruction of large residential spaces for embodied AI, offering a practical solution that avoids the drift and high image count of pinhole-camera pipelines.
CasaMaestro takes 20-50 sparse multi-view panoramas to directly predict metric depth and camera poses, enabling fast, drift-free 3D reconstruction of entire houses. It achieves high-quality results in both real and synthetic scenes.
The rise of home-deployed embodied AI systems is driving a growing need for fast, metric 3D reconstruction of residential spaces to support navigation, interaction, and long-horizon task execution. However, the commonly used pinhole-camera 3D reconstruction pipelines struggle to model large indoor residences efficiently due to their limited field of view, to which achieving full coverage across multiple rooms often requires thousands of images and incurs drift from long chains of incremental alignment. In this work, we present CasaMaestro (Spanish words meaning ``house'' and ``master''), a feedforward model that can take only twenty to fifty sparse multi-view indoor panoramas as input and directly predicts metric depth along with camera poses, allowing fast point-cloud reconstruction of the entire house with full coverage. CasaMaestro is the first model that supports house-scale reconstruction with multi-view panoramas. Experiments show that CasaMaestro can robustly provide high quality results in both real-world and synthetic scenes, which can serve as a strong foundation for acquiring house-scale 3D indoor assets to be applied in close-loop simulation.