Open-Set Source Tracing as Compositional Factors via Structured Prototypes
For researchers in synthetic speech detection, this work provides a more nuanced approach to source tracing that handles novel combinations of generative factors, improving robustness in open-set scenarios.
The paper redefines source tracing for synthetic speech as identifying compositional factors (architecture, training data, etc.) and proposes a framework using structured orthonormal prototypes with subspace partitioning. The method significantly outperforms angular-margin baselines in few-shot open-set identification on MLAAD.
Recent research expands beyond binary anti-spoofing with the emergence of Source Tracing, the task of identifying the specific generative origins of synthetic speech. However, current research often equates a "source" with its generative architecture. We propose redefining a source as a compositional tuple of Architecture, Training Data, and other training factors affecting the generated speech. We propose a framework using Structured Orthonormal Prototypes to minimize class overlap and intra-class variance. Our Subspace Partitioning strategy splits the embedding into architecture and data subspaces, while a residual subspace captures stochastic variability, enabling "compositional generalization" for novel factor combinations. This approach improves performance for partially seen sources and maintains robustness in fully open-set scenarios. MLAAD evaluations for Few-Shot open-set Identification show our approach significantly outperforms angular-margin baselines.