Model Releases
The Role of Disfluencies in Speech Translation
arXiv:2608.02138v1 Announce Type: new Abstract: Current speech translation systems, including SpeechLLMs, are trained on cleaned text and tend to strip disfluencies like filled pauses and false starts
arXiv:2608.02138v1 Announce Type: new Abstract: Current speech translation systems, including SpeechLLMs, are trained on cleaned text and tend to strip disfluencies like filled pauses and false starts rather than translate them. We show this comes at a cost: disfluencies carry meaning that gets lost when speech is cleaned up. To study this systematically, we introduce Uh-Mazing, a benchmark of human-translated, disfluency-annotated Switchboard speech covering English into eight target languages. Across these languages and several architectures, we find that false starts and self-repairs, not filled pauses or discourse markers, drive most of the translation-quality loss, and that models which fail to preserve a disfluency tend to omit it rather than mistranslate it. We show inference-time decoding can mitigate this without retraining, and release the benchmark and code.
Related
- MCAT: Scaling Many-to-Many Speech-to-Text Translation with MLLMs to 70 Languages
- F-Actor: Controllable Conversational Behaviour in Full-Duplex Models
- SpeechLLM Meets Federated Learning for End-to-End ASR: English and Italian Case Studies
Source: arXiv cs.CL | 2026-08-04