SDAIJun 15

Dual-Granularity Orthogonal Disentanglement for Generalizable Audio Deepfake Detection

arXiv:2606.165325.7
Predicted impact top 72% in SD · last 90 daysOriginality Incremental advance
AI Analysis

For audio deepfake detection, this work provides a method to improve cross-dataset generalization without complex architectures or adversarial training.

The paper tackles the problem of audio deepfake detectors failing to generalize across speakers due to implicit identity leakage. The proposed dual-granularity orthogonal disentanglement framework achieves 1.35%, 7.88%, and 21.58% EER on ASVspoof 2019 LA, ASVspoof 2021 DF, and In-the-Wild datasets, surpassing gradient reversal by 2.60% absolute on cross-dataset transfer.

Audio deepfake detectors often fail to generalize across speakers, as they learn speaker-identity features rather than synthesis artifacts, known as implicit identity leakage. Existing methods address this but incur architectural complexity or training instability. This paper proposes a dual-granularity orthogonal disentanglement framework enforcing feature independence at two levels: sample-level cosine orthogonality captures directional decorrelation, while batch-level cross-covariance regularization eliminates linear correlations across embedding dimensions. A curriculum disentanglement schedule progressively strengthens the orthogonality constraint without auxiliary networks or adversarial dynamics. Experiments on ASVspoof 2019 LA, ASVspoof 2021 DF, and In-the-Wild datasets demonstrate that the proposed method achieves 1.35%, 7.88%, and 21.58% equal error rates (EER), respectively, surpassing gradient reversal disentanglement by 2.60% absolute on cross-dataset transfer.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes