T2AV-Compass: Towards Unified Evaluation for Text-to-Audio-Video Generation
arXiv:2512.21094v2 Announce Type: replace Abstract: Text-to-Audio-Video (T2AV) generation aims to synthesize temporally coherent video and semantically synchronized audio from natural language, yet it