Cinematic Compositing Using Character-Environment-Harmonized Video Generation Models
For filmmakers and VFX artists, this provides an end-to-end solution for realistic compositing without expensive rendering, though it is domain-specific to cinematic compositing.
The paper tackles cinematic compositing by jointly modeling character-environment interactions (physical and lighting) using a video diffusion framework, achieving significant improvements over existing methods in dynamic video compositing quality.
Cinematic compositing aims to integrate green-screen characters into novel environments while maintaining physical and photometric realism. Previous methods often fail to capture the complex bidirectional interactions between characters and their surroundings, which we characterize as Character-to-Environment (C2E) physical interaction and Environment-to-Character (E2C) lighting harmonization. To address this, we propose an end-to-end video diffusion framework that jointly models C2E and E2C interactions, specifically handling the challenges of interactive props. Our approach introduces a tri-mask-guided architecture with RGB-D joint denoising to ensure physically consistent interactions among the character, props, and environment. We further develop an efficient prior-driven data curation pipeline to construct high-quality relighting pairs without expensive rendering. Finally, a reference-conditioned mechanism enables controllable environment synthesis and precise prop replacement. Extensive experiments demonstrate that our framework significantly outperforms existing methods in cinematic-quality dynamic video compositing.