From World Models to World Action Models: A Concise Tutorial for Robotics
For researchers and practitioners in embodied AI and robotics, this tutorial provides a structured taxonomy to navigate and compare world model approaches, though it is a conceptual survey rather than a novel contribution.
This tutorial clarifies the ambiguous scope of world models in robotics by presenting a design-space view and introducing world action models. It categorizes existing methods into observation-space and state-space models, compares their trade-offs, and summarizes four paradigms for connecting predicted futures with robot actions.
World models are increasingly used in embodied intelligence and generative simulation, yet their scope remains ambiguous across communities. This tutorial presents a design-space view of world models as action-conditioned predictive models that estimate the future evolution of task-relevant observations or states. We categorize existing methods into observation-space and state-space world models, comparing their trade-offs in visual fidelity, spatial structure, physical interpretability, and control usability. We further introduce world action models, which connect predicted futures with executable robot actions, and summarize four representative paradigms: imagine-then-execute, video-feature-conditioned action prediction, joint video-action modeling, and auxiliary video prediction for policy learning. The goal of this tutorial is to clarify the conceptual scope of world (action) models and provide a structured taxonomy for embodied prediction and control.