ASSDJun 5

Assessing True Generalisability of Audio-Visual Speech Recognisers

arXiv:2606.072598.0
Predicted impact top 35% in AS · last 90 daysOriginality Incremental advance
AI Analysis

Exposes overfitting in AVSR models for the speech recognition community, showing that current benchmarks are unreliable.

Current AVSR models achieve near-perfect performance on LRS3 but fail to generalize to a matched unseen test set, with universal performance collapse across five architectures. Audio-visual performance even lags behind audio-only settings.

Current Audio-Visual Speech Recognition (AVSR) models achieve near-perfect performance on the standard LRS3 benchmark, raising concerns of adaptive overfitting. To systematically assess true generalisability, we construct a highly controlled, unseen evaluation set subsampled from the massive MultiVSR dataset. Unlike standard out-of-distribution benchmarks, our subset strictly matches the acoustic, visual, and demographic distributions of the LRS3 test set. Evaluating five state-of-the-art architectures reveals a universal performance collapse, proving that current systems fail to generalise even under strictly aligned conditions. Through a fine-grained attribute analysis across seven factors, we isolate the specific drivers of this degradation. Furthermore, we uncover a profound lexical bias, expose distinct error patterns, and surprisingly reveal that audio-visual performance even lags behind audio-only settings. We release our matched test set for future benchmarking.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes