MMBench-Live: A Continuously Evolving Benchmark for Multimodal Models
DGX agentarXiv:2607.01813v1 Announce Type: cross Abstract: Evaluation benchmarks are essential for assessing vision-language models (VLMs), but most multimodal benchmarks are static, making them vulnerable to