CVJul 22

Vera: Identity-Faithful Human Subject-to-Video Generation

arXiv:2607.2024715.8
Predicted impact top 12% in CV · last 90 daysOriginality Incremental advance
AI Analysis

For researchers and practitioners in human-centric video generation, Vera provides a unified framework that significantly reduces identity drift and subject confusion in both single- and multi-person settings, addressing a critical bottleneck in the field.

Vera addresses identity drift in human subject-to-video generation, especially in multi-person scenarios, by introducing a million-pair identity-aligned dataset and two complementary designs (IFMS and RALA). It achieves improved identity consistency and reduced confusion compared to prior methods.

Subject-to-video (S2V) generation has made substantial progress in preserving reference subjects across diverse categories, yet generic subject consistency remains insufficient for human-centric generation. A video may appear globally consistent while identity-critical human details still drift across frames, poses, and interactions. This issue becomes more severe in multi-person scenarios, where incorrect identity-role binding leads to subject confusion, attribute swapping, and excessive copying of reference-specific appearance cues. We propose Vera, a unified human-centric S2V framework for single- and multi-person generation. We first construct a million-pair identity-aligned human image-video dataset through person-level cross-clip retrieval, providing explicit identity correspondence and diverse references. Built on this dataset, Vera introduces two complementary designs. Identity-Focal Masked Supervision (IFMS) strengthens identity-aware learning with spatially focused supervision while reducing interference from irrelevant artifacts. Reference-Aware Layer-wise Attention (RALA) regulates how video tokens interact with reference identity cues in the DiT backbone, preserving stable identity anchors and enhancing layer-aware identity readout. Extensive experiments demonstrate that Vera improves human identity consistency, multi-person subject binding, and motion naturalness, while reducing identity confusion and excessive reference-image copying.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes