CVAug 28, 2025

EmoCAST: Emotional Talking Portrait via Emotive Text Description

Yiguo Jiang, Xiaodong Cun, Yong Zhang, Yudian Zheng, Fan Tang, Chi-Man Pun

arXiv:2508.20615v16.21 citationsh-index: 36Has Code

Originality Incremental advance

AI Analysis

This addresses the need for more expressive and controllable talking head synthesis for applications like virtual avatars and entertainment, though it is incremental as it builds on existing diffusion methods.

The paper tackles the problem of generating emotional talking head videos with flexible control and natural motion by proposing EmoCAST, a diffusion-based framework that integrates text-driven emotional synthesis and audio-emotion interplay, achieving state-of-the-art performance in realism and synchronization.

Emotional talking head synthesis aims to generate talking portrait videos with vivid expressions. Existing methods still exhibit limitations in control flexibility, motion naturalness, and expression quality. Moreover, currently available datasets are primarily collected in lab settings, further exacerbating these shortcomings. Consequently, these limitations substantially hinder practical applications in real-world scenarios. To address these challenges, we propose EmoCAST, a diffusion-based framework with two key modules for precise text-driven emotional synthesis. In appearance modeling, emotional prompts are integrated through a text-guided decoupled emotive module, enhancing the spatial knowledge to improve emotion comprehension. To improve the relationship between audio and emotion, we introduce an emotive audio attention module to capture the interplay between controlled emotion and driving audio, generating emotion-aware features to guide more precise facial motion synthesis. Additionally, we construct an emotional talking head dataset with comprehensive emotive text descriptions to optimize the framework's performance. Based on the proposed dataset, we propose an emotion-aware sampling training strategy and a progressive functional training strategy that further improve the model's ability to capture nuanced expressive features and achieve accurate lip-synchronization. Overall, EmoCAST achieves state-of-the-art performance in generating realistic, emotionally expressive, and audio-synchronized talking-head videos. Project Page: https://github.com/GVCLab/EmoCAST

View on arXiv PDF Code

Similar