LG IT MLJul 5, 2022

Linear Jamming Bandits: Sample-Efficient Learning for Non-Coherent Digital Jamming

arXiv:2207.02365v11.85 citationsh-index: 45

Originality Incremental advance

AI Analysis

This work addresses the slow convergence issue in non-coherent digital jamming for communication security applications, representing an incremental improvement over existing bandit-based approaches.

The paper tackles the sample efficiency problem in learning optimal jamming strategies against digital modulation schemes by introducing a linear bandit algorithm that leverages action similarities and context features, resulting in significantly improved convergence behavior compared to prior methods.

It has been shown (Amuru et al. 2015) that online learning algorithms can be effectively used to select optimal physical layer parameters for jamming against digital modulation schemes without a priori knowledge of the victim's transmission strategy. However, this learning problem involves solving a multi-armed bandit problem with a mixed action space that can grow very large. As a result, convergence to the optimal jamming strategy can be slow, especially when the victim and jammer's symbols are not perfectly synchronized. In this work, we remedy the sample efficiency issues by introducing a linear bandit algorithm that accounts for inherent similarities between actions. Further, we propose context features which are well-suited for the statistical features of the non-coherent jamming problem and demonstrate significantly improved convergence behavior compared to the prior art. Additionally, we show how prior knowledge about the victim's transmissions can be seamlessly integrated into the learning framework. We finally discuss limitations in the asymptotic regime.

View on arXiv PDF

Similar