ASSDJun 15

Decoding while Adapting: Zero-Shot Online Speaker Adaptation via Audio-Textual Prompts for Elderly Speech Recognition

arXiv:2606.1653910.1
Predicted impact top 37% in AS · last 90 daysOriginality Incremental advance
AI Analysis

It addresses the need for real-time, speaker-adaptive ASR for elderly users, a domain with high variability and limited data.

This paper introduces a zero-shot online speaker adaptation method for elderly speech recognition using cross-utterance audio-textual prompts, achieving statistically significant WER/CER reductions of 0.61% and 1.22% absolute (2.99% and 4.48% relative) on two elderly speech datasets, with up to 9.83x speed-up over offline adaptation.

This paper proposes a novel cross-utterance audio-textual prompts based speaker adaptation approach for elderly speech recognition. It enables zero-shot, real-time adaptation to unseen speakers. Speech and text embeddings are extracted from the current and a few preceding utterances, before being fused in a cross-modal manner to produce compact speaker prompts that are more consistent than i/x-vectors and ECAPA-TDNN features. Experiments on the English DementiaBank Pitt and Cantonese JCCOCC MoCA elderly speech datasets suggest that the proposed online adaptation outperforms the speaker-independent (SI) model by statistically significant word error rate (WER) or character error rate (CER) reductions of 0.61% and 1.22% absolute (2.99% and 4.48% relative). Real-time factor (RTF) speed-up ratios of up to 9.83 times are obtained over offline batch-mode adaptation.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes