SDAICRMMJun 10

Feature-Aligned Speech Watermarking for Robustness to Reconstruction Distortions

arXiv:2606.11828v110.5
Predicted impact top 33% in SD · last 90 daysOriginality Incremental advance
AI Analysis

For audio watermarking applications, this method addresses the robustness-fidelity trade-off, improving robustness to reconstruction distortions without sacrificing perceptual quality.

The paper proposes a feature-aligned audio watermarking method that aligns watermarks with speech feature distributions, enabling higher watermark energy for robustness against speech reconstruction models while preserving imperceptibility. Experiments show comparable imperceptibility to existing methods and substantially improved robustness under both seen and unseen reconstruction models.

Audio watermarking aims to embed identifiable information into audio while remaining imperceptible. Existing methods adopt high-fidelity, low-energy designs to preserve perceptual quality, but the resulting watermarks lack robustness under suppression by speech reconstruction models. Improving robustness is challenging due to the inherent robustness-fidelity trade-off in existing designs, where increasing watermark energy improves robustness but reduces fidelity. To address this problem, we propose a feature-aligned watermarking method that aligns the watermark with the original speech feature distribution, allowing higher watermark energy to improve robustness while preserving imperceptibility. We use a pretrained speech codec to generate a pseudo-speech watermark and fuse it into the spectrogram of the input audio, with VAD loss and perceptual losses guiding embedding within voiced regions. Experiments show that our method maintains imperceptibility comparable to existing approaches while substantially improving robustness under both seen and unseen speech reconstruction models.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes