SDLGMMJun 15

Beyond Artifacts: Towards Generalizable Synthetic Song Detection via Music-Intrinsic Features

arXiv:2606.1661210.4
Predicted impact top 34% in SD · last 90 daysOriginality Incremental advance
AI Analysis

For researchers and practitioners in AI-generated audio detection, this work provides a more robust, generator-agnostic method that outperforms existing approaches on a challenging new benchmark.

The paper tackles synthetic song detection (SSD) by proposing Sofia, a framework using music-intrinsic features and a Mixture-of-Experts module, achieving an 18.5-point F1 improvement over the strongest baseline on the MUSIC8K-O benchmark.

The rapid advancement of AI music generators highlights the urgent need for reliable Synthetic Song Detection (SSD). Existing SSD methods often rely on low-level artifacts or fixed feature assumptions, struggling to capture generator-agnostic cues. To address this, we propose Sofia (Synthetic-song detection framework via music features), a flexible framework that models music-intrinsic attributes via feature-specific experts and an adaptive Mixture-of-Experts (MoE) module. By configuring Sofia with representative Vocal, Audio-effect, Global structure features, and their combinations, we present their individual and complementary contributions. To comprehensively evaluate our framework, we further construct MUSIC8K, a challenging benchmark featuring lastest emerging generators and realistic audio perturbations. Experiments show that Sofia learns generator-agnostic representations from music-intrinsic features, improving the F1 score by 18.5 points over the strongest baseline on MUSIC8K-O while maintaining strong robustness.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes