SynSFX: Multi-Model Sound Effects Synthesis Dataset for Deepfake Detection and Evaluation
This dataset provides a resource for studying and improving deepfake detection specifically for sound effects, an area where existing detectors perform poorly.
The authors introduce SynSFX, a large-scale dataset of 43,374 audio clips (26,452 synthetic, 16,922 real) from 7 text-to-audio models, to address the limited generalization of deepfake detectors to synthetic sound effects.
While audio deepfake detection has advanced significantly, representative detectors show limited generalization to synthetic sound effects. Existing environmental audio datasets such as EnvSDD provide important initial resources, but remain limited in scale and generation provenance for studying isolated sound-effect deepfakes. To support this direction, we present SynSFX, a large-scale corpus of 43374 clips (26452 synthetic, 16922 real) spanning 7 popular text-to-audio models.