Hierarchical Codec Diffusion for Video-to-Speech Generation
DGX agentarXiv:2604.15923v1 Announce Type: cross Abstract: Video-to-Speech (VTS) generation aims to synthesize speech from a silent video without auditory signals. However, existing VTS methods disregard the h