Adaptive and Robust Watermark for Generative Tabular Data
It provides a theoretically grounded watermarking method for tabular data, addressing the need for robust protection against misuse of synthetic data.
This paper proposes a watermarking algorithm for generative tabular data with theoretical guarantees on fidelity, detectability, robustness, and decoding hardness, validated on synthetic and real-world datasets.
In recent years, watermarking generative tabular data has become a prominent framework to protect against the misuse of synthetic data. However, while most prior work in watermarking methods for tabular data demonstrate a wide variety of desirable properties (e.g., high fidelity, detectability, robustness), the findings often emphasize empirical guarantees against common oblivious and adversarial attacks. In this paper, we study a flexible and robust watermarking algorithm for generative tabular data. Specifically, we demonstrate theoretical guarantees on the performance of the algorithm on metrics like fidelity, detectability, robustness, and hardness of decoding. The proof techniques introduced in this work may be of independent interest and may find applicability in other areas of machine learning. Finally, we validate our theoretical findings on synthetic and real-world tabular datasets.