Massive Open-Vocabulary Keyword Spotting

arXiv:2606.11279v15.3h-index: 2
Predicted impact top 78% in AS · last 90 daysOriginality Incremental advance
AI Analysis

This work addresses the scalability bottleneck in contextual biasing for keyword spotting, enabling processing of massive glossaries for speech recognition systems.

The paper tackles the problem of open-vocabulary keyword spotting for rare terms, proposing a system that reduces memory footprint by up to 128x while maintaining comparable entity recall without fine-tuning, even for unseen languages.

Automatic speech recognition systems have been shown to under-perform when it comes to transcribing words rarely seen in the training data, namely specialized terminology. Open-vocabulary keyword spotting, combined with contextual biasing, has been shown to mitigate this issue. However, existing systems can only handle glossaries of a few hundred terms without becoming an infeasible bottleneck. We propose a system that stores features with a memory footprint up to 128 times smaller than a comparable baseline and allows users to process massive databases while remaining open-vocabulary. Without fine-tuning the speech recognition model, our system achieves a comparable entity recall as uncompressed solutions, even in languages not seen during training.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes