DynaVieW: Schema-Guided World Modeling for Understanding Hierarchical Visual Dynamics
This work addresses the challenge of modeling temporal visual dynamics for multimodal LLMs, which is crucial for video understanding and simulation applications.
DynaVieW introduces a schema-guided world model that improves multimodal LLMs' ability to predict and simulate hierarchical visual dynamics in videos, achieving better consistency and controllability in visual narrative creation and world simulation tasks.
Multimodal LLMs struggle to systematically model the temporal evolution of visual scenes in videos or multi-image sequences. Such inputs require models to predict or simulate multiple levels of dynamic constituents, such as actions taken in the visual sequence, and the associated changes to the visual environment that result. To address this challenge, we propose a dynamic schema-guided world model, DynaVieW, optimized for visual dynamic prediction and simulation. DynaVieW achieves an in-depth understanding of visual dynamics by learning interleaved state-transition sequences, where states cover broad visual scenes from video keyframes, and transitions capture comprehensive dynamic constituents within a hierarchical schema. DynaVieW jointly models transition prediction and state simulation under a mixture-of-experts architecture, with a cross-expert selective attention and a schema token re-weighted loss, to ensure effective and robust learning. DynaVieW's understanding of visual dynamics boosts its downstream performance in visual narrative creation and world simulation, showing improved consistency, controllability, and instruction-following.