SDJun 10

The Hidden Cost of Pairwise Verification in Synthetic Speech Source Tracing

arXiv:2606.11666v110.3h-index: 7
Predicted impact top 35% in SD · last 90 daysOriginality Synthesis-oriented
AI Analysis

For researchers in synthetic speech detection and source attribution, this work identifies a fundamental limitation of pairwise metric-learning objectives in open-set verification tasks.

The paper shows that global anchoring outperforms pairwise verification for synthetic speech source tracing, achieving 8.61% EER versus 12-15% EER, and attributes the gap to pairwise objectives concentrating variance into fewer embedding directions.

Open-set source tracing is increasingly framed as a verification problem, motivating the use of pairwise metric-learning objectives from biometrics. We thus compare global anchoring and pairwise verification under matched backbones and a fixed data and epoch budget on MLAAD (in-domain) and STOPA (out-of-domain). In our runs, global anchoring yields lower in-domain error (8.61% EER) than pairwise variants (12-15% EER), even with rival mining and XLS-R finetuning. Because pairwise objectives optimize similarity directly, they concentrate variance into fewer embedding directions, reducing resolution among closely related generators. To test if this drives the drop, we impose a similar bottleneck to the globally supervised baseline, yet the baseline remains competitive. Together with an embedding-space analysis ($k_{99}$), these results suggest that the gap is not explained by dimensionality alone, but rather by the pairwise objective's shaping of the retained directions.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes