Research
Language corpora for the Dutch medical domain
arXiv:2604.25374v1 Announce Type: new Abstract: extbf{Background:} Dutch medical corpora are scarce, limiting NLP development. extbf{Methods:} We translated English datasets, identified medical text i
arXiv:2604.25374v1 Announce Type: new Abstract: extbf{Background:} Dutch medical corpora are scarce, limiting NLP development. extbf{Methods:} We translated English datasets, identified medical text in generic corpora, and extracted open Dutch medical resources. extbf{Results:} The resulting corpus comprises pm 35 billion tokens across the medical domain in about 100 million documents, freely available on Hugging Face. extbf{Conclusion:} This work establishes the first large-scale Dutch medical language corpus for pre-training and downstream NLP tasks.
Related
- BhashaSutra: A Task-Centric Unified Survey of Indian NLP Datasets, Corpora, and Resources
- English is Not All You Need: Systematically Exploring the Role of Multilinguality in LLM Post-Training
- Translate or Simplify First: An Analysis of Cross-lingual Text Simplification in English and French
Source: arXiv cs.CL | 2026-04-29