ASSDJun 13

DuraMark: Duration-Embedded Watermarking in LLM-based TTS

arXiv:2606.152649.9
Predicted impact top 40% in AS · last 90 daysOriginality Incremental advance
AI Analysis

Addresses the vulnerability of speech watermarking to generative attacks in deepfake detection for TTS systems.

DuraMark proposes a watermarking framework for LLM-based TTS that embeds information via syllable duration editing, achieving superior robustness against generative attacks compared to signal-level methods.

Large language model (LLM)-based text-to-speech (TTS) models have achieved remarkable voice cloning capabilities, raising concerns about potential deepfake misuse. Speech watermarking mitigates this by embedding traceable information into generated speech. Mainstream watermarking methods operate at the signal level (waveform or spectrogram), rendering the watermark vulnerable to generative attacks (e.g., neural codec and vocoder). To address this, we propose DuraMark, a robust information-level watermarking framework. It utilizes syllable duration editing to achieve watermark embedding. Specifically, DuraMark integrates a duration-controllable LLM-based TTS model to edit syllable durations during synthesis, coupled with a duration extractor to extract these durations for detection. Experiments demonstrate DuraMark's superior robustness against generative attacks, significantly outperforming signal-level baselines. Audio samples are available at https://muzw.github.io/duramark_demo/.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes