JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing
DGX agentarXiv:2512.13677v2 Announce Type: replace Abstract: In this paper, we present JoVA, a streamlined framework that unifies joint video-audio generation and editing. While existing methods often rely on