Applications
Low-Latency Turn-Taking via Context-Aware Preface Generation in a Real-World Dialogue Robot
arXiv:2607.23204v1 Announce Type: cross Abstract: Large language model (LLM)-based dialogue systems suffer response delays because generation begins only after final speech recognition. While fixed fi
arXiv:2607.23204v1 Announce Type: cross Abstract: Large language model (LLM)-based dialogue systems suffer response delays because generation begins only after final speech recognition. While fixed fillers are a workaround, they become unnatural over time. We propose a two-stage incremental framework that decouples prefatory-response preparation from speech onset. Once user intent becomes predictable, an intent readiness detector triggers LLM-based generation of a short prefatory response. Concurrently, a voice activity projection (VAP) model determines when to deliver it. Through a field experiment with a route-guidance robot in a shopping mall, we evaluated three conditions: no-filler, fixed-filler, and contextual-preface. Both fixed-filler and contextual-preface significantly reduced initial response latency relative to no-filler. Relative to fixed-filler, contextual-preface had significantly longer initial response latency but a significantly shorter initial-to-main gap. Exploratory ratings showed no significant differences. These results indicate a timing trade-off.
Related
- RAG-DIVE: A Dynamic Approach for Multi-Turn Dialogue Evaluation in Retrieval-Augmented Generation
- When Silence Matters: The Impact of Irrelevant Audio on Text Reasoning in Large Audio-Language Models
- From Black Box to Glass Box: Cross-Model ASR Disagreement to Prioto Review in Ambient AI Scribe Documentation
- Core-based Hierarchies for Efficient GraphRAG
- Omni-Embed-Audio: Leveraging Multimodal LLMs for Robust Audio-Text Retrieval
Source: arXiv cs.CL | 2026-07-28