The Hidden Cost of Pairwise Verification in Synthetic Speech Source Tracing
For researchers in synthetic speech detection and source attribution, this work identifies a fundamental limitation of pairwise metric-learning objectives in open-set verification tasks.
The paper shows that global anchoring outperforms pairwise verification for synthetic speech source tracing, achieving 8.61% EER versus 12-15% EER, and attributes the gap to pairwise objectives concentrating variance into fewer embedding directions.
Open-set source tracing is increasingly framed as a verification problem, motivating the use of pairwise metric-learning objectives from biometrics. We thus compare global anchoring and pairwise verification under matched backbones and a fixed data and epoch budget on MLAAD (in-domain) and STOPA (out-of-domain). In our runs, global anchoring yields lower in-domain error (8.61% EER) than pairwise variants (12-15% EER), even with rival mining and XLS-R finetuning. Because pairwise objectives optimize similarity directly, they concentrate variance into fewer embedding directions, reducing resolution among closely related generators. To test if this drives the drop, we impose a similar bottleneck to the globally supervised baseline, yet the baseline remains competitive. Together with an embedding-space analysis ($k_{99}$), these results suggest that the gap is not explained by dimensionality alone, but rather by the pairwise objective's shaping of the retained directions.