Walking in the Implicit: Interactive World Exploration via Neural Scene Representation
For interactive world exploration tasks, this paper proposes a new paradigm that improves long-horizon consistency and efficiency without relying on pretrained video backbones or 3D reconstructors.
This work introduces a scene-centric paradigm for interactive video generation that replaces frame-by-frame latent rollout with a fixed-length Neural Implicit Scene (NIS), achieving strong long-horizon consistency and favorable inference efficiency on public posed-view data.
Interactive video generation systems for camera-controlled world exploration roll out growing sequences of latent video frames, entangling state transition with high-frequency observation synthesis. We propose Walking in the Implicit, a scene-centric paradigm that changes the rollout variable from frame latents to a fixed-length, renderable implicit state, termed Neural Implicit Scene (NIS). This factorizes interactive generation into stochastic transition of a compact scene state and deterministic pose-conditioned rendering given the sampled state. We instantiate this paradigm as NeuWorld: a transformer VAE learns locally anchored NIS from sparse posed frames, and a diffusion transformer evolves NIS conditioned on future camera trajectories and geometry-aware retrieved history. By reusing the VAE encoder as a unified conditioner, NeuWorld maps camera, reference-image, and history cues into the same NIS modality, avoiding external heterogeneous encoders. Trained from scratch on public posed-view data without pretrained video backbones or auxiliary 3D reconstructors, NeuWorld achieves strong long-horizon consistency with favorable inference efficiency.