SI CL CYJun 9, 2020

EPIC30M: An Epidemics Corpus Of Over 30 Million Relevant Tweets

Junhua Liu, Trisha Singhal, Lucienne T. M. Blessing, Kristin L. Wood, Kwan Hui Lim

arXiv:2006.08369v25.911 citationsHas Code

Originality Synthesis-oriented

AI Analysis

This provides a new benchmark dataset for researchers in epidemiology and related fields to study multiple epidemics, though it is incremental as it builds on existing corpus efforts.

The authors tackled the lack of large-scale datasets for cross-epidemic analysis by presenting EPIC30M, a corpus of over 30 million tweets from 2006 to 2020, including subsets for diseases like Ebola, Cholera, and Swine Flu, with statistics and use cases demonstrating its utility.

Since the start of COVID-19, several relevant corpora from various sources are presented in the literature that contain millions of data points. While these corpora are valuable in supporting many analyses on this specific pandemic, researchers require additional benchmark corpora that contain other epidemics to facilitate cross-epidemic pattern recognition and trend analysis tasks. During our other efforts on COVID-19 related work, we discover very little disease related corpora in the literature that are sizable and rich enough to support such cross-epidemic analysis tasks. In this paper, we present EPIC30M, a large-scale epidemic corpus that contains 30 millions micro-blog posts, i.e., tweets crawled from Twitter, from year 2006 to 2020. EPIC30M contains a subset of 26.2 millions tweets related to three general diseases, namely Ebola, Cholera and Swine Flu, and another subset of 4.7 millions tweets of six global epidemic outbreaks, including 2009 H1N1 Swine Flu, 2010 Haiti Cholera, 2012 Middle-East Respiratory Syndrome (MERS), 2013 West African Ebola, 2016 Yemen Cholera and 2018 Kivu Ebola. Furthermore, we explore and discuss the properties of the corpus with statistics of key terms and hashtags and trends analysis for each subset. Finally, we demonstrate the value and impact that EPIC30M could create through a discussion of multiple use cases of cross-epidemic research topics that attract growing interest in recent years. These use cases span multiple research areas, such as epidemiological modeling, pattern recognition, natural language understanding and economical modeling.

View on arXiv PDF Code

Similar