ASSDJun 19

Bridging the Age Gap: Towards Detecting Neural Audio Codec Synthesized Elderly Speech Deepfake

arXiv:2606.2173511.3
Predicted impact top 27% in AS · last 90 daysOriginality Incremental advance
AI Analysis

Addresses the vulnerability of deepfake detection for elderly speech, a specific demographic often overlooked in existing benchmarks.

The paper introduces the Elderly CodecFake Detection (ECFD) task and the Elderly-CodecFake dataset, showing that existing detectors fail on elderly speech. The proposed BONSAI framework, fusing multimodal foundation models, achieves an average EER of 1.66%, outperforming baselines.

In this study, we introduce the Elderly CodecFake Detection (ECFD) task and release the Elderly-CodecFake (ECF) dataset in English and Chinese. We show that state-of-the-art CF detectors trained on previous benchmark CF datasets generalize poorly to elderly speech, revealing a critical vulnerability. We further hypothesize and demonstrate that multimodal foundation models (FMs) such as LanguageBind (LB) and ImageBind (IB) are more effective for ECFD due to their exposure to elderly content during cross-modal pretraining. Motivated by prior evidence that fusion of FMs enhances downstream performance, we explore fusion of FMs for ECFD. To this end, we propose BONSAI, a novel framework that employs Jensen-Shannon Divergence as the fusion mechanism. BONSAI with the fusion of LB and IB achieves an average EER (%) of 1.66 and outperforms individual FMs as well as competitive SOTA baselines, establishing a new benchmark for the ECFD task.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes