CVAIROJun 18

TriMotion: Modality-Agnostic Camera Control for Video Generation

arXiv:2606.2077417.9
Predicted impact top 18% in CV · last 90 daysOriginality Incremental advance
AI Analysis

For video generation researchers and practitioners, TriMotion provides a unified solution to camera control across diverse input modalities, addressing the limitation of existing single-modality methods.

TriMotion introduces a modality-agnostic framework for camera-controlled video generation that maps video, pose, and text inputs into a shared motion embedding space, enabling accurate camera trajectory following across all modalities. Experiments show high-quality video generation with precise trajectory control, and the shared space enables applications like motion composition and cross-modal interpolation.

Camera motion control is essential for directing viewpoint changes in generative systems. However, existing methods typically condition the generation process on a single specific modality, such as explicit pose trajectories or reference videos, limiting their ability to support heterogeneous user inputs. To address this limitation, we present TriMotion, a modality-agnostic framework for camera-controlled video generation that maps video, pose, and text inputs, describing the same camera trajectory into a shared motion embedding space. Learning such a space requires synchronized supervision across modalities. Therefore, we build the Motion Triplet Dataset by extending a Multi-Cam Video Dataset with geometry-grounded motion descriptions derived from camera extrinsics. We further introduce a latent motion consistency objective that leverages the motion embedding space to encourage the generated video to follow the target camera trajectory directly in latent space, avoiding the cost of pixel-space decoding. Extensive experiments show that TriMotion generates high-quality videos that accurately follow the target camera trajectories across all three modalities. Beyond standard generation, the shared motion embedding space also enables flexible applications such as sequential motion composition and cross-modal motion interpolation.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes