AICVLGJun 29

FacePlex: Full-Duplex Joint Speech-Facial Motion Generation for Conversational Avatars

arXiv:2606.3014511.9
Predicted impact top 48% in AI · last 90 daysOriginality Highly original
AI Analysis

This work addresses the problem of generating both speech and facial motion simultaneously in real-time for conversational avatars, which was previously unsolved.

FacePlex introduces a unified streaming framework for real-time joint generation of speech and synchronized facial motion in conversational avatars, achieving stronger lip-sync quality and motion fidelity than audio-driven baselines.

Natural face-to-face conversation requires real-time speech generation together with synchronized facial motion. Existing systems only partially address this problem: speech-only full-duplex models can generate speech in real time but do not produce facial motion, while audio-driven facial motion models animate a face from already available audio rather than jointly generating speech and motion online. To bridge this gap, we first formalize full-duplex joint speech-facial motion generation, where speech tokens and facial motion tokens are produced together every step. Building on this formulation, we propose FacePlex, a unified streaming framework with two key components. First, Rolling Flow Matching adapts flow matching to online motion generation by committing new motion frames at each streaming step. Second, Rolling Cross-Attention couples the streaming audio queue with the motion queue, allowing speech and facial motion to condition each other as generation progresses. Through extensive experiments, ablation studies, and a user study, we show that FacePlex enables full-duplex joint speech-facial motion generation under online streaming constraints, while achieving stronger lip-sync quality and motion fidelity than audio-driven facial motion baselines.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes