CLAug 10, 2016

An assessment of orthographic similarity measures for several African languages

arXiv:1608.03065v13 citations
Originality Synthesis-oriented
AI Analysis

This work addresses the challenge of adapting NLP tools for under-resourced Southern African languages, though it is incremental as it focuses on evaluation rather than new method development.

The study assessed orthographic similarity measures across several African languages to determine cross-language adaptability of NLP tools, finding that language clusters based on orthography do not align with traditional Guthrie zones and that tools like NLTK require significant customization for these languages.

Natural Language Interfaces and tools such as spellcheckers and Web search in one's own language are known to be useful in ICT-mediated communication. Most languages in Southern Africa are under-resourced, however. Therefore, it would be very useful if both the generic and the few language-specific NLP tools could be reused or easily adapted across languages. This depends on the notion, and extent, of similarity between the languages. We assess this from the angle of orthography and corpora. Twelve versions of the Universal Declaration of Human Rights (UDHR) are examined, showing clusters of languages, and which are thus more or less amenable to cross-language adaptation of NLP tools, which do not match with Guthrie zones. To examine the generalisability of these results, we zoom in on isiZulu both quantitatively and qualitatively with four other corpora and texts in different genres. The results show that the UDHR is a typical text document orthographically. The results also provide insight into usability of typical measures such as lexical diversity and genre, and that the same statistic may mean different things in different documents. While NLTK for Python could be used for basic analyses of text, it, and similar NLP tools, will need considerable customization.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes