SynCity 3000: Bootstrapping Scene-Scale 3D Diffusion
This work addresses the problem of generating large-scale, coherent 3D scenes for computer graphics and simulation, but the reliance on synthetic data and fine-tuning makes it an incremental improvement over existing image-to-3D methods.
SynCity 3000 generates globally coherent 3D scenes with fine-grained layout control by adapting image-to-3D generators as convolutional operators, fine-tuned on synthetic scene data. It produces large, coherent, and detailed scenes across diverse prompts and layouts.
We present SynCity 3000, a framework for generating 3D scenes that are globally coherent while enabling fine-grained layout control. Building on the ability of current image-to-3D generators to produce complex 3D assets from a single image, we extend this capability to the scale of entire scenes by adapting the generator to be applicable as a convolutional operator. We achieve this by fine-tuning the model on scene-like data generated by a new synthetic data engine, which we propose to address the scarcity of 3D scene data for training. The convolutional generator is then applied to a dimetric image of the entire scene, generated from the user prompt, resulting in 3D scenes of arbitrary size and complexity. Across diverse prompts and layouts, SynCity 3000 produces large, coherent, and detailed scenes, addressing the shortcomings of prior approaches to 3D scene generation.