ASLGJun 18

Low-Burden Data Augmentation for Dysarthric ASR via Zero-Shot Voice Cloning

arXiv:2606.198238.0
Predicted impact top 55% in AS · last 90 daysOriginality Incremental advance
AI Analysis

For dysarthric ASR researchers, zero-shot cloning offers a scalable augmentation method that reduces the need for extensive speaker-specific recordings.

Zero-shot voice cloning was used to augment training data for dysarthric ASR, achieving 26.00% WER on TORGO (vs. 24.44% with real data) and 11.45% relative improvement on SAP-1102, demonstrating a low-burden alternative to costly data collection.

Automatic speech recognition remains unreliable for dysarthric speech due to data scarcity and high inter-speaker variability. While synthetic data can address these gaps, traditional methods often require extensive speaker-specific data, reintroducing the collection bottleneck. We investigate zero-shot voice cloning as a low-burden augmentation strategy, using Higgs Audio V2 to clone speakers in the TORGO dataset. We fine-tune (FT) Whisper-medium on cloned, real, and hybrid data and evaluate on held-out real speech. Compared to the zero-shot (31.62%), Clone FT achieved a competitive 26.00% WER, nearly matching the 24.44% and 25.12% seen with Real and Hybrid FT, respectively. Notably, Clone and Hybrid FT outperform Real FT for moderate-severe speakers. Clone FT achieves the best results (11.45% relative) in cross-corpus evaluation on the SAP-1102. These results suggest that zero-shot cloning provides scalable training data that circumvents the costly data collection bottleneck.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes