CVJul 9

VSRo-200: A Romanian Visual Speech Recognition Dataset for Studying Supervision and Multimodal Robustness

arXiv:2607.081123.1h-index: 3
Predicted impact top 89% in CV · last 90 daysOriginality Incremental advance
AI Analysis

This provides a new resource and benchmark for visual speech recognition in a low-resource language, enabling systematic study of supervision quality and multimodal robustness.

The paper introduces VSRo-200, the first large-scale Romanian visual speech recognition dataset with 200 hours of video, and establishes a benchmark showing that pseudo-labels enable scalability while human annotations yield better performance at fixed scales. Multimodal fusion improves robustness under noise, and representations transfer to isolated word recognition, outperforming prior results.

We introduce VSRo-200, the first large-scale dataset for visual speech recognition (lip reading) in Romanian, comprising 200 hours of real-world podcast videos. All samples are annotated with pseudo-labels generated by a fine-tuned Romanian ASR model, while a subset of 100 hours is additionally transcribed by humans, enabling controlled analysis of supervision quality under a unified framework. Building on this dataset, we establish a benchmark for visual speech recognition in low-resource settings. We systematically study the impact of supervision quality, showing that while human annotations provide better performance at fixed data scales, pseudo-labels enable continued improvements through scalability. We further evaluate robustness under domain shift using curated out-of-distribution (OOD) test sets, and analyze audio-visual speech recognition (AVSR) under noisy conditions, where multimodal fusion significantly improves robustness compared to audio-only models. Finally, we demonstrate that representations learned on VSRo-200 transfer effectively to the LRRo benchmark for isolated word recognition, substantially outperforming previously reported results. Overall, VSRo-200 provides a new testbed for studying supervision, domain generalization, and multimodal fusion in low-resource visual speech recognition.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes