SDCRJun 5

A Large-Scale Per-Speaker Analysis of Re-identification Risk in Speech Anonymization

arXiv:2606.072107.7
Predicted impact top 52% in SD · last 90 daysOriginality Synthesis-oriented
AI Analysis

For researchers and practitioners in speech privacy, this work highlights the inadequacy of average-case metrics and the need for attacker- and anonymizer-conditioned evaluation protocols.

This paper analyzes per-speaker re-identification risk in speech anonymization across nearly 5,000 speakers, showing that vulnerability varies greatly and depends on the interaction between attacker, anonymizer, and speech duration, challenging the idea of intrinsic privacy risks.

Speech anonymization is commonly evaluated using averagecase metrics such as the equal error rate, which can hide large disparities in re-identification risks across individuals. In this paper, we conduct a large-scale per-speaker privacy analysis using a linkability-based metric under a worst-case scenario. Nearly 5,000 speakers are evaluated across multiple anonymization systems, attacker architectures, and conversation lengths. While linkability scores are highly polarized at the speaker level, the sets of easy to re-identify and hard to re-identify speakers vary substantially across configurations. We show that no single factor explains speaker vulnerability. Instead, the re-identification risk emerges from the interaction between the attacker, the anonymizer, and the amount of available speech. These results challenge the notion of intrinsic speaker-level privacy risks and emphasize the need for evaluation protocols that are explicitly conditioned on the attacker and anonymizer.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes