ROJun 5

Dreaming when Necessary: Advancing World Action Models with Adaptive Multi-Modal Reasoning

arXiv:2606.0708914.8
Predicted impact top 5% in RO · last 90 daysOriginality Incremental advance
AI Analysis

This work addresses the lack of adaptive reasoning in World Action Models for embodied agents, enabling more efficient and effective long-horizon task execution.

AdaWAM introduces adaptive multimodal reasoning for World Action Models, dynamically switching between textual and visual reasoning based on context, achieving substantial improvements in inference efficiency and outperforming state-of-the-art embodied policies on simulated and real-world tasks.

World Action Models (WAMs) offer a promising approach to embodied intelligence, yet existing methods rely heavily on video prediction as action priors and lack adaptive multimodal reasoning, limiting their effectiveness on long-horizon, complex tasks. We observe that WAMs require different multimodal reasoning modes under different execution contexts: textual reasoning is essential during task transitions to guide high-level action prediction, while visual reasoning is critical during fine-grained manipulation for precise control. Motivated by this observation, we propose \textbf{AdaWAM}, a world action model with adaptive multimodal reasoning abilities. AdaWAM integrates a lightweight dynamic router that autonomously triggers textual or visual reasoning as needed during task execution. Experiments on both simulated and real-world embodied tasks show that AdaWAM substantially improves inference efficiency while outperforming state-of-the-art embodied policies. Codes and demos are available at: https://adawam.github.io/.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes