CLJul 9

COALA: Robust Contextualized Speech-augmented Language Modeling for ASR via Contrastive Regularizer and Biasing Score Estimation

arXiv:2607.0811714.7h-index: 9
Predicted impact top 53% in CL · last 90 daysOriginality Incremental advance
AI Analysis

This work addresses the problem of recognizing domain-specific entities in ASR for multi-entity scenarios, offering a robust solution to training collapse and context-window limitations.

COALA proposes a robust framework for contextual biasing in ASR that maps SLM latent representations into a discriminative space to quantify matching intensity between audio and candidate entities, achieving superior performance on LibriSpeech across various biasing list scales.

Contextual biasing seeks to integrate external knowledge into automatic speech recognition (ASR) systems to accurately recognize domain-specific entities. In this paper, we propose COALA (Contextualized ASR Leveraging Biasing Scoring), a robust framework designed to enhance speech-augmented language models (SLMs) in complex multi-entity scenarios. Considering the inherent context-window limitations of SLMs, identifying relevant target entities from a large-scale biasing list is crucial for effective recognition. To this end, COALA maps SLM latent representations into a specialized discriminative space to quantify the matching intensity between audio segments and candidate entities. Furthermore, we address the training collapse in prior study when handling multi-target utterances-where multiple rare words co-occur. Experimental results on the LibriSpeech benchmark demonstrate that COALA consistently achieves superior contextual biasing performance across various biasing list scales.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes