Model Releases
InteracVid: Building a Real Interactive Audio-Visual Response Dataset from Live-Chat Videos
arXiv:2608.01157v1 Announce Type: new Abstract: Large language models have made text the default medium for human--AI interaction, buttext alone cannot express the full range of responses required by
arXiv:2608.01157v1 Announce Type: new Abstract: Large language models have made text the default medium for human--AI interaction, buttext alone cannot express the full range of responses required by multimodal assistants,avatars, and embodied agents. While recent audio-video generative models can synthesizehigh-fidelity synchronized content, existing supervision is largely descriptive:models are trained to render captions rather than to produce audio-visual responsescaused by external user interactions. We introduce extbf{InteracVid}, the firstopen-source large-scale dataset that addresses this missing supervision, so that everysample couples a preceding audio-visual context and an external stimulus with the realinteractive response that follows. We design a metadata-aware pipeline that extractsinteractive clips from long, noisy livestreams, yielding over extbf{454K}context-query-response triplets from more than extbf{59K} livestream videos andspanning conversation-centered, object-centric, procedural, embodied, and screen-basedscenarios. A ten-rater human study confirms that the extracted interactions are causal,natural, and temporally complete for both genuine and reconstructed queries. On aheld-out benchmark of extbf{100} genuine live-chat queries, fine-tuning on InteracVidimproves both interaction planning and audio-video response generation, and anindependent human evaluation reproduces the system ranking and the conclusions obtainedwith our automatic judge. These results highlight interaction-structured data as acritical foundation for interactive multimodal generation.
Source: arXiv cs.CV | 2026-08-04