Face and Voice Cross-modal Association with Learning Convex Feature Embedding
DGX agentarXiv:2607.28129v1 Announce Type: new Abstract: Face-and-voice association learning is one of the most challenging tasks in deep learning. In this paper, we propose a simple but powerful cross-modal f