CLMar 9, 2022

Language Diversity: Visible to Humans, Exploitable by Machines

Gábor Bella, Erdenebileg Byambadorj, Yamini Chandrashekar, Khuyagbaatar Batsuren, Danish Ashgar Cheema, Fausto Giunchiglia

arXiv:2203.04723v132.0638 citationsh-index: 58

Originality Synthesis-oriented

AI Analysis

This provides a resource for linguists and AI researchers to analyze and utilize language diversity, though it is incremental as it builds on existing lexical database concepts.

The paper introduces the Universal Knowledge Core (UKC), a large multilingual lexical database covering over a thousand languages, designed to make language diversity visually understandable for humans and formally exploitable by machines, enabling exploration of cross-lingual phenomena like shared meanings and lexical gaps.

The Universal Knowledge Core (UKC) is a large multilingual lexical database with a focus on language diversity and covering over a thousand languages. The aim of the database, as well as its tools and data catalogue, is to make the somewhat abstract notion of diversity visually understandable for humans and formally exploitable by machines. The UKC website lets users explore millions of individual words and their meanings, but also phenomena of cross-lingual convergence and divergence, such as shared interlingual meanings, lexicon similarities, cognate clusters, or lexical gaps. The UKC LiveLanguage Catalogue, in turn, provides access to the underlying lexical data in a computer-processable form, ready to be reused in cross-lingual applications.

View on arXiv PDF

Similar