DLCLMar 18, 2020

A Corpus of Adpositional Supersenses for Mandarin Chinese

arXiv:2003.08437v153.1999 citations
Originality Synthesis-oriented
AI Analysis

This addresses a gap in resources for cross-linguistic semantic analysis and multilingual disambiguation systems, though it is incremental as it adapts an existing framework to a new language.

The paper tackles the lack of annotated corpora for adposition semantics in Mandarin Chinese by creating the first broadly annotated corpus for this purpose, achieving high inter-annotator agreement on a translation of The Little Prince and showing that supersense categories are well-suited to Chinese despite syntactic differences from English.

Adpositions are frequent markers of semantic relations, but they are highly ambiguous and vary significantly from language to language. Moreover, there is a dearth of annotated corpora for investigating the cross-linguistic variation of adposition semantics, or for building multilingual disambiguation systems. This paper presents a corpus in which all adpositions have been semantically annotated in Mandarin Chinese; to the best of our knowledge, this is the first Chinese corpus to be broadly annotated with adposition semantics. Our approach adapts a framework that defined a general set of supersenses according to ostensibly language-independent semantic criteria, though its development focused primarily on English prepositions (Schneider et al., 2018). We find that the supersense categories are well-suited to Chinese adpositions despite syntactic differences from English. On a Mandarin translation of The Little Prince, we achieve high inter-annotator agreement and analyze semantic correspondences of adposition tokens in bitext.

Code Implementations1 repo
Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes