Research
Testimole-Conversational: A 30-Billion-Word Italian Discussion Board Corpus (1996-2024) for Language Modeling and Sociolinguistic Research
arXiv:2602.14819v2 Announce Type: replace Abstract: We present 'Testimole-conversational' a massive collection of discussion boards messages in the Italian language. The large size of the corpus, more
arXiv:2602.14819v2 Announce Type: replace Abstract: We present "Testimole-conversational" a massive collection of discussion boards messages in the Italian language. The large size of the corpus, more than 30B word-tokens (1996-2024), renders it an ideal dataset for native Italian Large Language Models'pre-training. Furthermore, discussion boards' messages are a relevant resource for linguistic as well as sociological analysis. The corpus captures a rich variety of computer-mediated communication, offering insights into informal written Italian, discourse dynamics, and online social interaction in wide time span. Beyond its relevance for NLP applications such as language modelling, domain adaptation, and conversational analysis, it also support investigations of language variation and social phenomena in digital communication. The resource will be made freely available to the research community.
Related
- Current LLMs still cannot 'talk much' about grammar modules: Evidence from syntax
- Floating or Suggesting Ideas? A Large-Scale Contrastive Analysis of Metaphorical and Literal Verb-Object Constructions
- Splits! Flexible Sociocultural Linguistic Investigation at Scale
- Paragraph Segmentation Revisited: Towards a Standard Task for Structuring Speech
- Differentially Private Language Generation and Identification in the Limit
Source: arXiv cs.CL | 2026-04-10