CLAIJun 10

AfriSUD: A Dependency Treebank Collection for Evaluating Models on African Languages

arXiv:2606.12708v118.5
Predicted impact top 47% in CL · last 90 daysOriginality Incremental advance
AI Analysis

It provides a benchmark for evaluating NLP models on underrepresented African languages, highlighting limitations of current architectures.

The paper introduces AfriSUD, the first large-scale collection of syntactically annotated treebanks for nine African languages, and evaluates models on POS tagging and dependency parsing, finding a significant syntax gap where models fail to fully capture African-language syntax.

Despite their linguistic diversity and global significance, African languages remain underrepresented in research and resources to support NLP. We aim to bridge this gap by introducing AfriSUD, the first large-scale collection of syntactically annotated treebanks for nine diverse African languages spanning major language families and regions across Sub-Saharan Africa. Using the Surface-Syntactic Universal Dependencies (SUD) framework, our community-led effort provides high-quality, native-speaker verified data that capture typological key features such as agglutination and tone. We evaluate a range of models on AfriSUD for part-of-speech tagging and dependency parsing including non-transformer baselines, multilingual pretrained encoders, and LLMs. Our results reveal a significant syntax gap, where models still show clear limitations across the nine languages, suggesting that existing architectures may not fully capture the structural diversity of African-language syntax.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes