CVAISDJun 11, 2025

InterActHuman: Multi-Concept Human Animation with Layout-Aligned Audio Conditions

arXiv:2506.09984v118 citationsh-index: 12
Originality Highly original
AI Analysis

This addresses the need for controllable multi-concept human-centric videos in applications like animation and virtual reality, representing a novel method for a known bottleneck.

The paper tackles the problem of generating multi-concept human animations with precise per-identity control, discarding the single-entity assumption of prior methods, and achieves high-quality video generation validated by empirical results.

End-to-end human animation with rich multi-modal conditions, e.g., text, image and audio has achieved remarkable advancements in recent years. However, most existing methods could only animate a single subject and inject conditions in a global manner, ignoring scenarios that multiple concepts could appears in the same video with rich human-human interactions and human-object interactions. Such global assumption prevents precise and per-identity control of multiple concepts including humans and objects, therefore hinders applications. In this work, we discard the single-entity assumption and introduce a novel framework that enforces strong, region-specific binding of conditions from modalities to each identity's spatiotemporal footprint. Given reference images of multiple concepts, our method could automatically infer layout information by leveraging a mask predictor to match appearance cues between the denoised video and each reference appearance. Furthermore, we inject local audio condition into its corresponding region to ensure layout-aligned modality matching in a iterative manner. This design enables the high-quality generation of controllable multi-concept human-centric videos. Empirical results and ablation studies validate the effectiveness of our explicit layout control for multi-modal conditions compared to implicit counterparts and other existing methods.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes