CiteVQA: Benchmarking Evidence Attribution for Trustworthy Document Intelligence
arXiv:2605.12882v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) have significantly advanced document understanding, yet current Doc-VQA evaluations score only the final answ