CLHCJun 30

What Counts as an Error? Dual-Reference Benchmarking for Atypical ASR

arXiv:2606.3111212.3
Predicted impact top 67% in CL · last 90 daysOriginality Synthesis-oriented
AI Analysis

For researchers and practitioners evaluating ASR on atypical speech, this work highlights the need to choose the appropriate transcription reference to avoid biased model selection.

The paper reveals that ASR evaluations conflate verbatim and intended transcription references, leading to misleading performance rankings. Benchmarking 11 models on stuttered speech shows significant performance disparities depending on which reference is used.

ASR systems have been often reported to underperform on atypical speech. An often conflated compounding factor is the existence of two valid transcription references: verbatim (actual produced speech, including repetitions/prolongations) and intended (the canonical form of the text with disfluencies removed) in atypical speech recognition depending on context and use-case. Most ASR evaluations conflate this duality into a single ground truth and reward systems that delete disfluencies, ignoring verbatim faithfulness. We benchmark 11 ASR models from encoder-decoder, CTC and transducer families using both verbatim and intended references on atypical stuttered speech as a case study. Our quantitative assessment underlines the disparity in model performance and rankings using the two transcript styles. Through this analysis, we highlight the importance of selecting a suitable transcription reference for valid model selection depending on the use-case, particularly for atypical ASR.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes