Research
AfriSUD: A Dependency Treebank Collection for Evaluating Models on African Languages
arXiv:2606.12708v2 Announce Type: replace-cross Abstract: Despite their linguistic diversity and global significance, African languages remain underrepresented in research and resources to support NLP
arXiv:2606.12708v2 Announce Type: replace-cross Abstract: Despite their linguistic diversity and global significance, African languages remain underrepresented in research and resources to support NLP. We aim to bridge this gap by introducing AfriSUD, the first large-scale collection of syntactically annotated treebanks for nine diverse African languages spanning major language families and regions across Sub-Saharan Africa. Using the Surface-Syntactic Universal Dependencies (SUD) framework, our community-led effort provides high-quality, native-speaker verified data that capture typological key features such as agglutination and tone. We evaluate a range of models on AfriSUD for part-of-speech tagging and dependency parsing including non-transformer baselines, multilingual pretrained encoders, and LLMs. Our results reveal a significant syntax gap, where models still show clear limitations across the nine languages, suggesting that existing architectures may not fully capture the structural diversity of African-language syntax.
Related
- Opportunities and Challenges of Large Language Models for Low-Resource Languages in Humanities Research
- Evaluating the Effect of Linguistic Relatedness on Cross-Lingual Transfer in Large Multilingual Automatic Speech Recognition
- The Grammar of Transformers: A Systematic Review of Interpretability Research on Syntactic Knowledge in Language Models
Source: arXiv cs.AI | 2026-09-02