MMArch: Benchmarking Multimodal Reasoning Grounded in Architectural Evidence
DGX agentarXiv:2608.09281v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) perform strongly on engineering imagery, yet existing benchmarks mostly test drawing recognition, information e