FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision
arXiv:2608.01392v1 Announce Type: new Abstract: Text-conditioned human motion generation has made rapid progress with the emergence of large-scale motion--language datasets. However, even datasets wit