CL AI IR LG SISep 25, 2021

Overview of the CLEF-2019 CheckThat!: Automatic Identification and Verification of Claims

Tamer Elsayed, Preslav Nakov, Alberto Barrón-Cedeño, Maram Hasanain, Reem Suwaileh, Giovanni Da San Martino, Pepa Atanasova

arXiv:2109.15118v15.7122 citations

Originality Synthesis-oriented

AI Analysis

It addresses the problem of automating fact-checking for researchers and practitioners, but is incremental as it builds on a previous edition with expanded tasks and languages.

The paper presents the CLEF-2019 CheckThat! Lab, which tackled automatic identification and verification of claims in English and Arabic, with tasks including prioritizing claims for fact-checking and ranking web pages for usefulness, involving 47 registered teams and 14 submissions.

We present an overview of the second edition of the CheckThat! Lab at CLEF 2019. The lab featured two tasks in two different languages: English and Arabic. Task 1 (English) challenged the participating systems to predict which claims in a political debate or speech should be prioritized for fact-checking. Task 2 (Arabic) asked to (A) rank a given set of Web pages with respect to a check-worthy claim based on their usefulness for fact-checking that claim, (B) classify these same Web pages according to their degree of usefulness for fact-checking the target claim, (C) identify useful passages from these pages, and (D) use the useful pages to predict the claim's factuality. CheckThat! provided a full evaluation framework, consisting of data in English (derived from fact-checking sources) and Arabic (gathered and annotated from scratch) and evaluation based on mean average precision (MAP) and normalized discounted cumulative gain (nDCG) for ranking, and F1 for classification. A total of 47 teams registered to participate in this lab, and fourteen of them actually submitted runs (compared to nine last year). The evaluation results show that the most successful approaches to Task 1 used various neural networks and logistic regression. As for Task 2, learning-to-rank was used by the highest scoring runs for subtask A, while different classifiers were used in the other subtasks. We release to the research community all datasets from the lab as well as the evaluation scripts, which should enable further research in the important tasks of check-worthiness estimation and automatic claim verification.

View on arXiv PDF

Similar