MCIF: Multimodal Crosslingual Instruction-Following Benchmark from Scientific Talks
DGX agentarXiv:2507.19634v4 Announce Type: replace-cross Abstract: Recent advances in large language models have laid the foundation for multimodal LLMs (MLLMs), which unify text, speech, and vision within a s