CY AIMay 29, 2023

Game of Tones: Faculty detection of GPT-4 generated content in university assessments

Mike Perkins, Jasper Roe, Darius Postma, James McGaughran, Don Hickerson

arXiv:2305.18081v1154 citations

Originality Synthesis-oriented

AI Analysis

This addresses academic integrity concerns for universities by showing that current AI detection tools are vulnerable to evasion through prompt engineering, though it is incremental as it builds on existing research about AI in education.

This study tested whether university faculty could detect GPT-4 generated content in assessments with Turnitin's AI detection tool, finding that while the tool flagged 91% of AI submissions, it only identified 54.8% of the content as AI-generated, and faculty reported 54.5% of these submissions for misconduct, with AI-generated work scoring similarly (52.3) to genuine submissions (54.4).

This study explores the robustness of university assessments against the use of Open AI's Generative Pre-Trained Transformer 4 (GPT-4) generated content and evaluates the ability of academic staff to detect its use when supported by the Turnitin Artificial Intelligence (AI) detection tool. The research involved twenty-two GPT-4 generated submissions being created and included in the assessment process to be marked by fifteen different faculty members. The study reveals that although the detection tool identified 91% of the experimental submissions as containing some AI-generated content, the total detected content was only 54.8%. This suggests that the use of adversarial techniques regarding prompt engineering is an effective method in evading AI detection tools and highlights that improvements to AI detection software are needed. Using the Turnitin AI detect tool, faculty reported 54.5% of the experimental submissions to the academic misconduct process, suggesting the need for increased awareness and training into these tools. Genuine submissions received a mean score of 54.4, whereas AI-generated content scored 52.3, indicating the comparable performance of GPT-4 in real-life situations. Recommendations include adjusting assessment strategies to make them more resistant to the use of AI tools, using AI-inclusive assessment where possible, and providing comprehensive training programs for faculty and students. This research contributes to understanding the relationship between AI-generated content and academic assessment, urging further investigation to preserve academic integrity.

View on arXiv PDF

Similar