Adversarial Creation and Detection of AI-Generated Social Bot Content
For social media platforms and researchers combating misinformation, this work provides a more robust detection method for AI-generated bot content, though it is an incremental improvement over existing approaches.
The paper addresses the lack of ground-truth data for detecting AI-generated social bot content by proposing an adversarial methodology that models impersonation of real users, curating a multilingual, cross-platform dataset. Training on this data yields detection models that significantly outperform existing content-based bot detection methods on real-world, out-of-distribution data.
The convergence of large language models and social bots allows malicious actors to manipulate the information ecosystem by generating human-like content at scale. Existing models for detecting AI-generated content often fail in the wild, primarily due to the lack of ground-truth data. We address this gap through an adversarial methodology that models the impersonation of real social media users by malicious actors. Using this methodology, we curate a multilingual, cross-platform dataset of paired human and AI-generated messages. Training on such adversarial data yields accurate detection of AI-generated text. Our approach significantly outperforms existing models for content-based bot detection in real-world, out-of-distribution data.