CLDec 31, 2020

Open Korean Corpora: A Practical Report

Won Ik Cho, Sangwhan Moon, Youngsook Song

arXiv:2012.15621v231.0995 citationsHas Code

Originality Synthesis-oriented

AI Analysis

This work aims to improve resource availability and promote research for less-resourced languages like Korean by providing a curated list of open corpora.

This paper addresses the perception of Korean as a low-resource language by curating and reviewing existing open Korean corpora. It describes institution-level resource development and lists current open datasets for various tasks.

Korean is often referred to as a low-resource language in the research community. While this claim is partially true, it is also because the availability of resources is inadequately advertised and curated. This work curates and reviews a list of Korean corpora, first describing institution-level resource development, then further iterate through a list of current open datasets for different types of tasks. We then propose a direction on how open-source dataset construction and releases should be done for less-resourced languages to promote research.

View on arXiv PDF

Similar