DL IRFeb 11, 2020

Testing of Support Tools for Plagiarism Detection

Tomáš Foltýnek, Dita Dlabolová, Alla Anohina-Naumeca, Salim Razı, Július Kravjar, Laima Kamzola, Jean Guerrero-Dib, Özgür Çelik, Debora Weber-Wulff

arXiv:2002.04279v19.2102 citations

Originality Synthesis-oriented

AI Analysis

This work addresses the effectiveness of plagiarism detection tools for educators and researchers, highlighting their limitations as incremental improvements.

The paper tested 15 web-based text-matching systems for plagiarism detection across multiple languages and document types, finding that while some systems help identify plagiarized content, they often miss plagiarism and falsely flag non-plagiarized material.

There is a general belief that software must be able to easily do things that humans find difficult. Since finding sources for plagiarism in a text is not an easy task, there is a wide-spread expectation that it must be simple for software to determine if a text is plagiarized or not. Software cannot determine plagiarism, but it can work as a support tool for identifying some text similarity that may constitute plagiarism. But how well do the various systems work? This paper reports on a collaborative test of 15 web-based text-matching systems that can be used when plagiarism is suspected. It was conducted by researchers from seven countries using test material in eight different languages, evaluating the effectiveness of the systems on single-source and multi-source documents. A usability examination was also performed. The sobering results show that although some systems can indeed help identify some plagiarized content, they clearly do not find all plagiarism and at times also identify non-plagiarized material as problematic.

View on arXiv PDF

Similar