CLSep 25, 2025

RedHerring Attack: Testing the Reliability of Attack Detection

arXiv:2509.20691v14.91 citationsh-index: 4EMNLP

Originality Incremental advance

AI Analysis

This addresses a critical vulnerability in adversarial defense systems for NLP, exposing how adversaries can undermine detection reliability, though it is incremental in exploring a new threat model.

The paper tackles the reliability of attack detection models in NLP by proposing RedHerring, an attack that modifies text to cause detection models to incorrectly predict attacks while keeping classifiers accurate, resulting in detection accuracy drops of 20-71 points across datasets.

In response to adversarial text attacks, attack detection models have been proposed and shown to successfully identify text modified by adversaries. Attack detection models can be leveraged to provide an additional check for NLP models and give signals for human input. However, the reliability of these models has not yet been thoroughly explored. Thus, we propose and test a novel attack setting and attack, RedHerring. RedHerring aims to make attack detection models unreliable by modifying a text to cause the detection model to predict an attack, while keeping the classifier correct. This creates a tension between the classifier and detector. If a human sees that the detector is giving an ``incorrect'' prediction, but the classifier a correct one, then the human will see the detector as unreliable. We test this novel threat model on 4 datasets against 3 detectors defending 4 classifiers. We find that RedHerring is able to drop detection accuracy between 20 - 71 points, while maintaining (or improving) classifier accuracy. As an initial defense, we propose a simple confidence check which requires no retraining of the classifier or detector and increases detection accuracy greatly. This novel threat model offers new insights into how adversaries may target detection models.

View on arXiv PDF

Similar