Frédéric Agnès

CL
h-index4
3papers
101citations
Novelty23%
AI Score27

3 Papers

5.4CLMay 24, 2017Code
Deep Investigation of Cross-Language Plagiarism Detection Methods

Jeremy Ferrero, Laurent Besacier, Didier Schwab et al.

This paper is a deep investigation of cross-language plagiarism detection methods on a new recently introduced open dataset, which contains parallel and comparable collections of documents with multiple characteristics (different genres, languages and sizes of texts). We investigate cross-language plagiarism detection methods for 6 language pairs on 2 granularities of text units in order to draw robust conclusions on the best methods while deeply analyzing correlations across document styles and languages.

5.7CLApr 5, 2017Code
CompiLIG at SemEval-2017 Task 1: Cross-Language Plagiarism Detection Methods for Semantic Textual Similarity

Jeremy Ferrero, Frederic Agnes, Laurent Besacier et al.

We present our submitted systems for Semantic Textual Similarity (STS) Track 4 at SemEval-2017. Given a pair of Spanish-English sentences, each system must estimate their semantic similarity by a score between 0 and 5. In our submission, we use syntax-based, dictionary-based, context-based, and MT-based methods. We also combine these methods in unsupervised and supervised way. Our best run ranked 1st on track 4a with a correlation of 83.02% with human annotations.

2.1CLFeb 10, 2017Code
UsingWord Embedding for Cross-Language Plagiarism Detection

J. Ferrero, F. Agnes, L. Besacier et al.

This paper proposes to use distributed representation of words (word embeddings) in cross-language textual similarity detection. The main contributions of this paper are the following: (a) we introduce new cross-language similarity detection methods based on distributed representation of words; (b) we combine the different methods proposed to verify their complementarity and finally obtain an overall F1 score of 89.15% for English-French similarity detection at chunk level (88.5% at sentence level) on a very challenging corpus.