Research
MCGA: A Multi-task Classical Chinese Literary Genre Audio Corpus
arXiv:2601.09270v3 Announce Type: replace Abstract: With the rapid advancement of Multimodal Large Language Models (MLLMs), their potential has gained significant attention in Chinese Classical Studie
arXiv:2601.09270v3 Announce Type: replace Abstract: With the rapid advancement of Multimodal Large Language Models (MLLMs), their potential has gained significant attention in Chinese Classical Studies (CCS). While existing research primarily focuses on text and visual modalities, the audio corpus within this domain remains largely underexplored. To bridge this gap, we introduce the Multi-task Classical Chinese Literary Genre Audio Corpus (MCGA), a 119-hour corpus comprising 22,000 audio samples. It encompasses a diverse range of literary genres across six tasks: Automatic Speech Recognition (ASR), Speech-to-Text Translation (S2TT), Speech Emotion Captioning (SEC), Spoken Question Answering (SQA), Speech Understanding (SU), and Speech Reasoning (SR). Through the evaluation of ten MLLMs, our experimental results demonstrate that current MLLMs still face substantial challenges on the MCGA test set. Furthermore, we introduce a domain-specific metric for SEC and a metric to measure the consistency between speech and text capabilities. We release MCGA to the public to facilitate the development of more robust MLLMs. MCGA Corpus: https://github.com/yxduir/MCGA
Related
- LingoLoop Attack: Trapping MLLMs via Linguistic Context and State Entrapment into Endless Loops
- BiT-MCTS: A Theme-based Bidirectional MCTS Approach to Chinese Fiction Generation
- Testimole-Conversational: A 30-Billion-Word Italian Discussion Board Corpus (1996-2024) for Language Modeling and Sociolinguistic Research
- Current LLMs still cannot 'talk much' about grammar modules: Evidence from syntax
Source: arXiv cs.CL | 2026-04-14