CVROJul 10

Causally Debiased Latent Action Model for Embodied Action Conditioned World Models

arXiv:2607.0918514.9h-index: 6
Predicted impact top 18% in CV · last 90 daysOriginality Incremental advance
AI Analysis

For embodied AI researchers, CD-LAM provides a practical method to learn controllable world models from unlabeled video, reducing the need for costly action-labeled data.

Latent action models (LAMs) for action-conditioned world models suffer from action-irrelevant bias, entangling dynamics with visual factors like backgrounds. CD-LAM introduces three fine-tuning objectives that reduce this bias, achieving 12x fewer robot-action adaptation updates and improved controllability on 2B and 14B backbones.

Action-conditioned world models (ACWMs) aim to simulate future observations conditioned on embodied actions, offering a promising foundation for robot planning, policy evaluation, and data augmentation. However, learning controllable ACWMs requires large-scale action-labeled data, which remains costly to collect in the real world. Latent action models (LAMs) mitigate this bottleneck by inferring latent actions from unlabeled videos, but existing LAMs are typically trained with reconstruction-only objectives and therefore entangle action-relevant dynamics with action-irrelevant visual factors such as backgrounds and untouched objects. In this work, we identify this action-irrelevant bias as a key obstacle to controllable ACWMs and introduce evaluation metrics to measure latent-action bias, action following, and robustness. We propose CD-LAM, a causally debiased framework for LAM-based ACWMs. CD-LAM introduces three efficient fine-tuning objectives: embodiment-centric reconstruction, action-centric contrastive learning, and latent space calibration, which together encourage embodiment-focused, action-aware, and calibrated non-collapsed latent action representations. Experiments on 2B and 14B ACWM backbones show that CD-LAM substantially improves latent-action controllability, downstream robot-action following, visual fidelity, and adaptation efficiency, requiring only 6k fine-tuning steps and more than 12$\times$ fewer robot-action adaptation updates than the baseline.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes