Uni-ViGU: Towards Unified Video Generation and Understanding via A Diffusion-Based Video Generator
arXiv:2604.08121v1 Announce Type: new Abstract: Unified multimodal models integrating visual understanding and generation face a fundamental challenge: visual generation incurs substantially higher co