Safety

JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing

arXiv:2512.13677v2 Announce Type: replace Abstract: In this paper, we present JoVA, a streamlined framework that unifies joint video-audio generation and editing. While existing methods often rely on

DGX agentpaper
safetyarxiv-cs-cv

arXiv:2512.13677v2 Announce Type: replace Abstract: In this paper, we present JoVA, a streamlined framework that unifies joint video-audio generation and editing. While existing methods often rely on fragmented, task-specific architectures or complex fusion mechanisms, JoVA employs native joint representation learning for direct video, audio, and text interaction in a dual-branch architecture. This design eliminates redundant alignment modules and effectively unifies diverse multimodal tasks within a single model. Furthermore, we utilize channel-wise conditioning for flexible image and video reference to avoid massive token expansion, alongside a mouth-area loss to enhance lip alignment. To fully empower and systematically evaluate this framework, we construct a comprehensive training corpus encompassing video-audio generation and editing datasets, and introduce unified benchmarks tailored for these multimodal tasks. Extensive experiments demonstrate that JoVA achieves state-of-the-art performance across benchmarks, establishing it as an extensible framework for versatile content creation. Project page: https://visual-ai.github.io/jova

Source: arXiv cs.CV | 2026-08-03

Loading related sources…