SDAIASJun 22

HALAS: A Human-Annotated Dataset of Hallucinations of Modern ASR Systems

arXiv:2606.2304811.6
Predicted impact top 29% in SD · last 90 daysOriginality Incremental advance
AI Analysis

For ASR researchers, this provides the first realistic benchmark for detecting and mitigating hallucinations in modern ASR systems, revealing that current detection methods are inadequate.

The authors introduce HALAS, the first human-annotated dataset of naturally occurring hallucinations from seven state-of-the-art ASR models on real earnings call recordings. Their benchmark shows that proxy metrics achieve 81% ROC-AUC for hallucination detection, while state-of-the-art methods only reach 53.1% F1 score.

End-to-end Automatic Speech Recognition (ASR) systems hallucinate on natural speech, yet existing mitigation methods are typically evaluated on non-speech or artificially corrupted audio. We introduce HALAS, the first human-annotated dataset of naturally occurring hallucinations from seven state-of-the-art ASR models on real unprocessed earnings call recordings. HALAS provides span-level labels, enabling analysis of hallucination patterns and their severity. Our analysis reveals strong cross-model vocabulary overlap and confirms that hallucinations also occur for almost correctly transcribed speech (characterized by a low Word Error Rate). The proposed benchmark with HALAS shows that the character and semantic-level metrics used as a proxy for hallucination detection reach 81% ROC-AUC, while state-of-the-art detection methods achieve an F1 score of only 53.1%. As such, HALAS establishes the first rigorous non-artificial benchmark for the detection and mitigation of ASR hallucinations.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes