Research

Language corpora for the Dutch medical domain

arXiv:2604.25374v1 Announce Type: new Abstract: extbf{Background:} Dutch medical corpora are scarce, limiting NLP development. extbf{Methods:} We translated English datasets, identified medical text i

DGX agentpaper
researcharxiv-cs-cl

arXiv:2604.25374v1 Announce Type: new Abstract: extbf{Background:} Dutch medical corpora are scarce, limiting NLP development. extbf{Methods:} We translated English datasets, identified medical text in generic corpora, and extracted open Dutch medical resources. extbf{Results:} The resulting corpus comprises pm 35 billion tokens across the medical domain in about 100 million documents, freely available on Hugging Face. extbf{Conclusion:} This work establishes the first large-scale Dutch medical language corpus for pre-training and downstream NLP tasks.

Related

Source: arXiv cs.CL | 2026-04-29

Loading related sources…