LGCRMLJul 10

Statistically Undetectable Backdoors in Deep Neural Networks

arXiv:2607.095327.6
Predicted impact top 42% in LG · last 90 daysOriginality Highly original
AI Analysis

This work reveals a fundamental power asymmetry between model trainers and users, showing that backdoors can be planted without leaving any statistical trace, which has serious implications for the security and trustworthiness of deep learning models.

The authors demonstrate that an adversarial model trainer can plant backdoors in deep neural networks that are statistically undetectable in the white-box setting, with the backdoored and honest models being close in total variation distance. The backdoor enables invariance-based adversarial examples for every input, while without it, generating such examples is provably impossible under standard cryptographic assumptions.

We show how an adversarial model trainer can plant backdoors in a large class of deep, feedforward neural networks. These backdoors are statistically undetectable in the white-box setting, meaning that the backdoored and honestly trained models are close in total variation distance, even given the full descriptions of the models (e.g., all of the weights). The backdoor provides access to invariance-based adversarial examples for every input, mapping distant inputs to unusually close outputs. However, without the backdoor, it is provably impossible (under standard cryptographic assumptions) to generate any such adversarial examples in polynomial time. Our theoretical and preliminary empirical findings demonstrate a fundamental power asymmetry between model trainers and model users.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes