ASSDJun 19

Towards Detecting Neural Audio Codec Synthesized Heart Sounds

arXiv:2606.217277.4
Predicted impact top 60% in AS · last 90 daysOriginality Incremental advance
AI Analysis

This work addresses the need for detecting synthetic heart sounds generated by neural audio codecs, which is important for ensuring the integrity of medical audio recordings.

The paper introduces the task of detecting neural audio codec synthesized heart sounds (SHAC) and releases the first benchmark dataset, CARDIOFAKE. The proposed GROOT fusion framework achieves state-of-the-art performance by combining MFCC and WavLM features.

In this paper, we introduce Synthetic Heart Sound Detection (SHAC), a task aimed at identifying phonocardiograms (PCGs) synthesized using neural audio codecs (NACs). To facilitate research in this direction, we release CARDIOFAKE, the first benchmark dataset for SHAC containing both real and codec-synthesized PCGs. We benchmark spectral representations (MFCC, LFCC) and self-supervised learning (SSL) representations (e.g., WavLM) for the task. Furthermore, we propose GROOT, a fusion framework that integrates spectral and SSL features for leveraging their complementary behavior. Experiments show that GROOT, combining MFCC and WavLM, achieves state-of-the-art performance, outperforming individual representations and competitive baselines.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes