CVJun 21

Customizing Video Portraits via Identity-ActionDecoupling

arXiv:2606.2234712.5
Predicted impact top 36% in CV · last 90 daysOriginality Incremental advance
AI Analysis

For researchers in text-to-video generation, this work addresses a key bottleneck in controlling facial dynamics while preserving identity, offering a more expressive and controllable solution.

The paper tackles identity-preserving text-to-video generation, where prior methods produce monotonous or inaccurate facial movements. The proposed Identity-Action Decoupling (IaD) framework, with two novel loss functions, generates videos with high identity consistency and rich, prompt-aligned expressions without subject-specific fine-tuning.

Identity-Preserving Text-to-Video Generation (IPT2V) seeks to synthesize a temporally coherent video from a reference image and a textual description, while simultaneously preserving the subject's identity and allowing fine-grained control over facial dynamics. Although recent methods such as ID-Animator and ConsisID inject identity features only at inference time, they ignored the ID-irrelevant information contained in Facial embedding, leading to monotonous or inaccurate facial movements that poorly follow the prompt. We introduce Identity-Action Decoupling (IaD) framework as well as two loss function Identity Decoupling Loss and Text Alignment Loss to solve this problem. Without any subject-specific fine-tuning, IaD yields videos that (1) maintain cross-temporal identity consistency and (2) exhibit rich, controllable expressions and scene variations that closely match the input text.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes