Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement
This work addresses the need for better cross-modal fusion in audio-visual speech enhancement, offering a simple yet effective improvement for robust speech recovery in noisy environments.
The paper proposes augmenting diffusion-based unsupervised audio-visual speech enhancement with a contrastive audio-visual loss to improve cross-modal alignment, achieving consistent gains in interference suppression, signal reconstruction, and perceptual quality, especially at low SNRs.
Audio-visual speech enhancement (AVSE) exploits visual cues such as lip movements to recover speech in noisy environments. Recent work introduced diffusion-based unsupervised AVSE, where a speech diffusion model conditioned on visual features via cross-attention is trained and used as a data-driven prior for posterior sampling-based speech enhancement. Despite promising performance over its audio-only counterpart, the impact of explicitly enforcing cross-modal alignment in the fusion remains unclear. In this work, we propose to augment the diffusion training objective with a contrastive audio-visual loss to encourage stronger use of visual information while keeping the posterior sampling framework unchanged. Experiments across matched and mismatched test data show consistent improvements in interference suppression, signal reconstruction, and perceptual quality, with the largest gains at low SNRs. Code is available at https://github.com/ cexauce/AV-CA-DiffUSE