Research
Computational Narrative Understanding for Expressive Text-to-Speech
arXiv:2509.04072v2 Announce Type: replace-cross Abstract: Recent advances in text-to-speech (TTS) have been driven by large, multi-domain speech corpora, yet the expressive potential of audiobook data
arXiv:2509.04072v2 Announce Type: replace-cross Abstract: Recent advances in text-to-speech (TTS) have been driven by large, multi-domain speech corpora, yet the expressive potential of audiobook data remains underexamined. We argue that human-narrated audiobooks, particularly fictional works, contain rich and diverse prosodic cues arising from the natural alternation between neutral narration and expressive character dialogue. Building from this observation, we introduce LibriQuote, a large-scale 5.3K hours of expressive speech drawn from character quotations. Each quote is supplemented with contextual pseudo-labels for speech verbs and adverbs that characterize the intended delivery of direct speech (e.g., "he whispered softly"). We found that fine-tuning a flow-matching model on LibriQuote yields substantial improvements in expressivity and intelligibility, while training from scratch enhances expressiveness of an autoregressive TTS model. Benchmarking on LibriQuote-test highlights significant variability across systems in generating expressive speech. We publicly release the dataset, code, and evaluation resources to facilitate reproducibility. Audio samples can be found at https://libriquote.github.io/.
Related
- TMD-TTS: A Unified Tibetan Multi-Dialect Text-to-Speech Framework for U-Tsang, Amdo and Kham Speech Dataset Generation
- iPhoneme: Brain-to-Text Communication for ALS Using ConformerXL Decoding
- Cross-lingual Matryoshka Representation Learning across Speech and Text
- From Speech-to-Spatial: Grounding Utterances on A Live Shared View with Augmented Reality
Source: arXiv cs.CL | 2026-04-22