CVAILGJul 31

Have I Seen You? Embedding Behavior Signals Synthetic Face Dataset Membership

arXiv:2607.291442.6h-index: 5
Predicted impact top 91% in CV · last 90 daysOriginality Synthesis-oriented
AI Analysis

For privacy researchers and practitioners using synthetic data, this work demonstrates a concrete privacy risk that synthetic data can retain traces of real training data, highlighting the need for stronger leakage mitigation.

The paper investigates whether synthetic face datasets leak information about the real datasets used to train their generators. They propose a dataset-level membership inference attack and find that it can identify the synthetic training dataset in 100% of cases and the generator's source dataset in 54.5% of cases across a wide range of models and datasets.

Synthetic face datasets are increasingly used to reduce privacy exposure and data access constraints in biometric recognition. Yet the generators that produce these datasets are trained on real faces, so synthetic data may still reveal their real source data. We study this risk through a dataset-level membership inference attack that first identifies the synthetic dataset used to train a face recognizer and then infers the real dataset used to train the generator. Across 11 face recognition models, 11 synthetic datasets, and 7 real datasets, the attack recovers the synthetic training dataset in 100% of cases and identifies the generator's source dataset in 54.5% of cases. These results show that synthetic data can retain dataset-level traces of real training data and that privacy-preserving deployment requires stronger leakage mitigation.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes