MATCHA: Matching Text via Contrastive Semantic Alignment
DGX agentarXiv:2605.27345v1 Announce Type: new Abstract: Reliable evaluation is essential for understanding large language model (LLM) performance, yet today's go-to metrics, namely token-overlap scores (e.g.,