Wild3R: Feed-Forward 3D Gaussian Splatting from Unconstrained Sparse Photo Collection
For 3D reconstruction from real-world photo collections, Wild3R removes the need for per-scene optimization while handling challenging conditions, though it is incremental over existing feed-forward 3DGS.
Wild3R tackles feed-forward 3D Gaussian Splatting from unconstrained sparse photo collections with diverse lighting and transient objects. It introduces the WildCity dataset (200 scenes, 337,500 images) and achieves results competitive with per-scene optimization methods, outperforming existing feed-forward approaches.
Feed-forward 3D Gaussian Splatting (3DGS) removes the need for time-consuming per-scene optimization required by traditional 3DGS. However, existing feed-forward approaches struggle with real-world photo collections that include diverse lighting conditions and transient objects. In this paper, we present Wild3R, a feed-forward approach for unconstrained sparse photo collections. The main bottleneck is the lack of training data that provides multiple viewpoints, a variety of illuminations, and transient variations necessary for learning robust scene representations. To address this, we introduce the WildCity dataset, which comprises 200 scenes, 170 lighting conditions, and transient objects, resulting in 337,500 images in total. By leveraging the dataset, our model learns appearance consistency across viewpoints conditioned on reference views, while removing transient content. Extensive experiments demonstrate that our method outperforms existing feed-forward approaches and achieves results competitive with prior per-scene optimization-based methods.