Research

AfriSUD: A Dependency Treebank Collection for Evaluating Models on African Languages

arXiv:2606.12708v2 Announce Type: replace-cross Abstract: Despite their linguistic diversity and global significance, African languages remain underrepresented in research and resources to support NLP

DGX agentpaper
researcharxiv-cs-ai

arXiv:2606.12708v2 Announce Type: replace-cross Abstract: Despite their linguistic diversity and global significance, African languages remain underrepresented in research and resources to support NLP. We aim to bridge this gap by introducing AfriSUD, the first large-scale collection of syntactically annotated treebanks for nine diverse African languages spanning major language families and regions across Sub-Saharan Africa. Using the Surface-Syntactic Universal Dependencies (SUD) framework, our community-led effort provides high-quality, native-speaker verified data that capture typological key features such as agglutination and tone. We evaluate a range of models on AfriSUD for part-of-speech tagging and dependency parsing including non-transformer baselines, multilingual pretrained encoders, and LLMs. Our results reveal a significant syntax gap, where models still show clear limitations across the nine languages, suggesting that existing architectures may not fully capture the structural diversity of African-language syntax.

Related

Source: arXiv cs.AI | 2026-09-02

Loading related sources…