CLJun 10, 2022

Building an Icelandic Entity Linking Corpus

arXiv:2206.05014v1583 citationsh-index: 17
Originality Synthesis-oriented
AI Analysis

This addresses the lack of entity linking resources for Icelandic, which is incremental as it applies existing methods to a new language.

The paper tackles the problem of creating the first Entity Linking corpus for Icelandic by comparing a combined method using a multilingual model (mGENRE) with Wikipedia API Search (WAPIS) to WAPIS alone, achieving 53.9% coverage versus 30.9%.

In this paper, we present the first Entity Linking corpus for Icelandic. We describe our approach of using a multilingual entity linking model (mGENRE) in combination with Wikipedia API Search (WAPIS) to label our data and compare it to an approach using WAPIS only. We find that our combined method reaches 53.9% coverage on our corpus, compared to 30.9% using only WAPIS. We analyze our results and explain the value of using a multilingual system when working with Icelandic. Additionally, we analyze the data that remain unlabeled, identify patterns and discuss why they may be more difficult to annotate.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes