Research
RUMLEM: A Dictionary-Based Lemmatizer for Romansh
arXiv:2604.11233v1 Announce Type: new Abstract: Lemmatization -- the task of mapping an inflected word form to its dictionary form -- is a crucial component of many NLP applications. In this paper, we
arXiv:2604.11233v1 Announce Type: new Abstract: Lemmatization -- the task of mapping an inflected word form to its dictionary form -- is a crucial component of many NLP applications. In this paper, we present RUMLEM, a lemmatizer that covers the five main varieties of Romansh as well as the supra-regional standard variety Rumantsch Grischun. It is based on comprehensive, community-driven morphological databases for Romansh, enabling RUMLEM to cover 77-84% of the words in a typical Romansh text. Since there is a dedicated database for each Romansh variety, an additional application of RUMLEM is variety-aware language classification. Evaluation on 30'000 Romansh texts of varying lengths shows that RUMLEM correctly identifies the variety in 95% of cases. In addition, a proof of concept demonstrates the feasibility of Romansh vs. non-Romansh language classification based on the lemmatizer.
Related
- GLeMM: A large-scale multilingual dataset for morphological research
- MIXAR: Scaling Autoregressive Pixel-based Language Models to Multiple Languages and Scripts
- Paragraph Segmentation Revisited: Towards a Standard Task for Structuring Speech
- Testimole-Conversational: A 30-Billion-Word Italian Discussion Board Corpus (1996-2024) for Language Modeling and Sociolinguistic Research
- HistLens: Mapping Idea Change across Concepts and Corpora
Source: arXiv cs.CL | 2026-04-14