Impact Analysis of Speech Representation Learning Models for Acoustic Side-Channel Attack
For researchers in acoustic side-channel attacks, this work provides a new benchmark dataset and demonstrates that KAN-based fine-tuning can improve attack performance, though the gains are incremental and domain-specific.
This paper introduces KEYAC, a dataset for acoustic side-channel attacks, and evaluates six speech representation learning models, finding that partial fine-tuning improves performance but models struggle with VoIP codec generalization. Using Kolmogorov-Arnold Networks for fine-tuning achieves new state-of-the-art results on KEYAC.
Acoustic side-channel attacks (ASCA) on keyboards have gained increasing attention, yet impact of speech representation learning models in ASCA remains unexplored. Addressing this, we introduce KEYAC, a dataset designed to analyze representation generalization for ASCA under both standard and VoIP codec settings. On KEYAC, we evaluate six representation learning models under zero-shot and partial fine-tuning settings using fully connected and convolutional networks. Results show that while partial fine-tuning improves performance, models struggle to generalize across VoIP codecs. We hypothesize this limitation stems from inadequate modeling of nonlinear feature interactions in conventional fine-tuning architectures. To address this, we employ Kolmogorov-Arnold Networks (KAN) for fine-tuning. Empirical results show that KAN-based fine-tuning consistently outperforms the baselines and establishes a new state-of-the-art on KEYAC.