Safety
In-Sync: Adaptation of Speech Aware Large Language Models for ASR with Word Level Timestamp Predictions
arXiv:2604.22817v1 Announce Type: cross Abstract: Recent advances in speech-aware language models have coupled strong acoustic encoders with large language models, enabling systems that move beyond tr
arXiv:2604.22817v1 Announce Type: cross Abstract: Recent advances in speech-aware language models have coupled strong acoustic encoders with large language models, enabling systems that move beyond transcription to produce richer outputs. Among these, word-level timestamp prediction is critical for applications such as captioning, media search, and multimodal synchronization, yet it is often handled by external alignment tools. In this work, we extend an existing speech-aware language model to predict timestamps directly alongside transcripts. We introduce a set of novel lightweight training strategies that improve alignment robustness while preserving recognition quality. Experiments across multiple datasets show that these strategies not only enhance timestamp accuracy, but also yield gains in overall ASR performance. Together, they demonstrate an efficient and unified approach to speech recognition with precise timestamp prediction.
Related
- Beyond Transcription: Unified Audio Schema for Perception-Aware AudioLLMs
- End-to-end Contrastive Language-Speech Pretraining Model For Long-form Spoken Question Answering
- VIBE: Voice-Induced open-ended Bias Evaluation for Large Audio-Language Models via Real-World Speech
- Identifying and typifying demographic unfairness in phoneme-level embeddings of self-supervised speech recognition models
- Character Beyond Speech: Leveraging Role-Playing Evaluation in Audio Large Language Models via Reinforcement Learning
Source: arXiv cs.CL | 2026-04-28