ASCVIVJun 5, 2019

Investigating the Lombard Effect Influence on End-to-End Audio-Visual Speech Recognition

arXiv:1906.02112v45.911 citations
Originality Incremental advance
AI Analysis

It addresses robustness in speech recognition for noisy environments, but is incremental as it focuses on a previously ignored factor.

This paper tackles the problem of audio-visual speech recognition in noisy environments by investigating the impact of the Lombard effect, showing that adding Lombard speech to training significantly improves performance in real scenarios.

Several audio-visual speech recognition models have been recently proposed which aim to improve the robustness over audio-only models in the presence of noise. However, almost all of them ignore the impact of the Lombard effect, i.e., the change in speaking style in noisy environments which aims to make speech more intelligible and affects both the acoustic characteristics of speech and the lip movements. In this paper, we investigate the impact of the Lombard effect in audio-visual speech recognition. To the best of our knowledge, this is the first work which does so using end-to-end deep architectures and presents results on unseen speakers. Our results show that properly modelling Lombard speech is always beneficial. Even if a relatively small amount of Lombard speech is added to the training set then the performance in a real scenario, where noisy Lombard speech is present, can be significantly improved. We also show that the standard approach followed in the literature, where a model is trained and tested on noisy plain speech, provides a correct estimate of the video-only performance and slightly underestimates the audio-visual performance. In case of audio-only approaches, performance is overestimated for SNRs higher than -3dB and underestimated for lower SNRs.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes