Detecting LLM-Generated Tokens in Human--LLM Coauthored Text
For researchers and practitioners needing to localize LLM-generated content in mixed-authorship documents, this method provides a simple, training-free approach with strong empirical performance.
This paper addresses the need for fine-grained detection of LLM-generated tokens in human-LLM coauthored text. The proposed token-level method, using adaptive smoothing of detection scores, achieves favorable mean square error performance and outperforms baselines on synthetic and realistic datasets.
The rise of human-AI collaborative writing has created a growing need for fine-grained detection methods that support localizing likely LLM-generated content in mixed-authorship documents. Existing methods for detecting LLM-generated text mainly focus on document-level classification and cannot identify which parts of the text are generated by LLMs. This paper introduces a new method to address this urgent need. Our method operates at the token level, the natural unit of modern language models, and builds on existing token-level detection scores. The key idea is to smooth adjacent token scores to reduce their variability, while using an adaptive Lepski-type rule to select the bandwidth according to the local authorship structure. Our method is simple to implement and does not require token-level labeled data for training. Theoretically, we characterize this trade-off and show that the proposed method achieves favorable mean square error performance in estimating the underlying signal. Empirically, we demonstrate strong performance of our method against a wide range of baselines in both synthetic datasets and a realistic dataset. We deploy a publicly accessible website that implements the methods as well.