SDCLJun 22

Synthesizing the Lombard Effect: Multi-Level Control of Speech Clarity and Vocal Effort in TTS

arXiv:2606.231764.7
Predicted impact top 78% in SD · last 90 daysOriginality Incremental advance
AI Analysis

This work addresses the need for controllable speech clarity in TTS for challenging listening environments, offering a novel method for simulating the Lombard effect.

The paper introduces a flow-matching TTS model with multi-level control of vocal effort and articulation, achieving continuous and disentangled control and word-level emphasis. Experiments show improved clarity features and intelligibility gains in noisy conditions, simulating the Lombard effect.

Humans tend to speak louder and clearer in challenging environments, such as noisy conditions or when addressing hearingimpaired listeners, which is called Lombard effect. To simulate this behavior in speech synthesis systems, we introduce a flow-matching based text-to-speech (TTS) model trained with vocal effort and articulation pseudo-labels. The proposed model achieves continuous and disentangled control of vocal effort and articulation, while also enabling word-level emphasis for clarifying specific segments of an utterance. Experimental results show that these control mechanisms effectively improve clarityrelated acoustic features. Furthermore, speech-in-noise experiments demonstrate that our model successfully simulates the intelligibility gains of human clear speech in noisy conditions.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes