MoDiCoL: A Modular Diagnostic Continual Learning Dataset for Robust Speech Recognition
For ASR researchers, this provides a controlled benchmark to analyze model robustness under realistic, evolving conditions, though it is an incremental contribution as a new dataset.
MoDiCoL introduces a modular diagnostic continual learning dataset for ASR to study robustness under co-occurring distribution shifts, and evaluates three continual learning strategies, showing how robustness is acquired, transferred, and forgotten.
Modern Automatic Speech Recognition (ASR) systems have made remarkable progress on standard benchmarks, yet performance gaps have emerged under real-world distribution shifts, caused by recording conditions, accents, speech impairments, and noise. Existing datasets and benchmarks typically isolate these factors, which overlooks their co-occurrence in real-world applications. In this paper, we argue that model robustness can be treated as a dynamic capability that continually develops, and we introduce MoDiCoL, a Modular Diagnostic Continual Learning dataset designed for controlled analysis of linguistic content, speaker characteristics, and acoustic environments. Furthermore, we propose a real-world-inspired continual learning curriculum to simulate incremental updates and study how robustness is acquired, transferred, and forgotten. We evaluate three continual learning strategies and provide detailed insights into robustness under evolving conditions.