SDAICRJun 18

Exploiting Neural Audio Codec Latents for Adversarial Audio Attacks

arXiv:2606.208934.6
Predicted impact top 80% in SD · last 90 daysOriginality Incremental advance
AI Analysis

This work addresses the need for efficient, real-time adversarial attacks on audio classification systems, which is important for security researchers and practitioners evaluating model robustness.

The authors propose a generative adversarial attack framework that operates in the latent space of a neural audio codec, achieving up to 99% targeted attack success rates with sub-7 ms inference, outperforming generative baselines and reducing latency by 24x.

Deep learning-based audio classification systems, including automatic speaker verification, are vulnerable to adversarial attacks. Realistic real-time threat assessment remains difficult because optimization-based methods, such as projected gradient descent (PGD) and Carlini-Wagner, require costly iterative updates in the high-dimensional waveform domain. Generative attacks allow single-shot synthesis but often introduce perceptible artifacts or depend on computationally intensive architectures, while diffusion and autoregressive approaches incur high inference latency. To address this gap, we propose a generative attack framework operating in the continuous latent space of a neural audio codec. A conditional generator synthesizes class-specific perturbations in a single forward pass and decodes them into adversarial waveforms. Our method achieves targeted attack success rates up to 99% with sub-7 ms inference, outperforming generative baselines while reducing latency by 24x.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes