ASAISDJun 23

DTT-BSR+: A Generative-Regression Cascade for Music Source Restoration

arXiv:2606.241276.2
Predicted impact top 66% in AS · last 90 daysOriginality Incremental advance
AI Analysis

For music source separation and restoration, this work provides a novel decoupling approach that outperforms existing systems, though improvements are incremental.

DTT-BSR+ introduces a two-stage cascade for music source restoration that separates distribution fitting from signal reconstruction, achieving state-of-the-art MMSNR on five stems and improving over single-stage DTT-BSR across all stems.

Music source restoration (MSR) requires jointly addressing source unmixing and the inversion of non-linear production effects. Current methods struggle to achieve accurate target signal reconstruction while maintaining semantic consistency. To address this limitation, we propose DTT-BSR+, a two-stage cascade MSR system that decouples distribution fitting from signal reconstruction into separate stages. A generative DTT-BSR separator in the first stage produces stems matching the prior of clean sources, and a modified Demucs network in the second stage enhances the first stage output using time-domain and multi-resolution spectral losses. DTT-BSR+ improves multi-mel signal-to-noise ratio (MMSNR) over the single-stage DTT-BSR across all stems, and surpasses the state-of-the-art X-LANCE MSR system on five stems. We also reveal through Fréchet Audio Distance (FAD) decomposition an implicit trade-off between signal reconstruction accuracy and semantic distribution fitting across stems.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes