CVAILGJan 24, 2025

Dreamweaver: Learning Compositional World Models from Pixels

NVIDIA
arXiv:2501.14174v55 citationsh-index: 8ICLR
Originality Highly original
AI Analysis

It addresses the challenge of replicating human-like compositional reasoning in AI for video understanding and generation, which is incremental as it builds on existing world modeling approaches.

The paper tackles the problem of learning compositional world models from raw videos to decompose perceptions into objects and attributes, enabling the generation of novel future simulations by recombining concepts, and demonstrates that the model outperforms state-of-the-art baselines in world modeling across multiple datasets.

Humans have an innate ability to decompose their perceptions of the world into objects and their attributes, such as colors, shapes, and movement patterns. This cognitive process enables us to imagine novel futures by recombining familiar concepts. However, replicating this ability in artificial intelligence systems has proven challenging, particularly when it comes to modeling videos into compositional concepts and generating unseen, recomposed futures without relying on auxiliary data, such as text, masks, or bounding boxes. In this paper, we propose Dreamweaver, a neural architecture designed to discover hierarchical and compositional representations from raw videos and generate compositional future simulations. Our approach leverages a novel Recurrent Block-Slot Unit (RBSU) to decompose videos into their constituent objects and attributes. In addition, Dreamweaver uses a multi-future-frame prediction objective to capture disentangled representations for dynamic concepts more effectively as well as static concepts. In experiments, we demonstrate our model outperforms current state-of-the-art baselines for world modeling when evaluated under the DCI framework across multiple datasets. Furthermore, we show how the modularized concept representations of our model enable compositional imagination, allowing the generation of novel videos by recombining attributes from previously seen objects. cun-bjy.github.io/dreamweaver-website

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes