CL AI LGFeb 14, 2024

Ten Words Only Still Help: Improving Black-Box AI-Generated Text Detection via Proxy-Guided Efficient Re-Sampling

Yuhui Shi, Qiang Sheng, Juan Cao, Hao Mi, Beizhe Hu, Danding Wang

arXiv:2402.09199v133 citationsh-index: 18IJCAI

Originality Incremental advance

AI Analysis

This addresses the societal issue of AI misuse in fake news and academic dishonesty by providing a more practical detection method for black-box scenarios, though it is incremental as it builds on existing re-sampling techniques.

The paper tackles the problem of detecting AI-generated text in black-box settings by estimating word generation probabilities via efficient re-sampling, achieving improved macro F1 scores across multiple datasets and settings while reducing computational costs.

With the rapidly increasing application of large language models (LLMs), their abuse has caused many undesirable societal problems such as fake news, academic dishonesty, and information pollution. This makes AI-generated text (AIGT) detection of great importance. Among existing methods, white-box methods are generally superior to black-box methods in terms of performance and generalizability, but they require access to LLMs' internal states and are not applicable to black-box settings. In this paper, we propose to estimate word generation probabilities as pseudo white-box features via multiple re-sampling to help improve AIGT detection under the black-box setting. Specifically, we design POGER, a proxy-guided efficient re-sampling method, which selects a small subset of representative words (e.g., 10 words) for performing multiple re-sampling in black-box AIGT detection. Experiments on datasets containing texts from humans and seven LLMs show that POGER outperforms all baselines in macro F1 under black-box, partial white-box, and out-of-distribution settings and maintains lower re-sampling costs than its existing counterparts.

View on arXiv PDF

Similar