Model Releases
CArtBench: Evaluating Vision-Language Models on Chinese Art Understanding, Interpretation, and Authenticity
arXiv:2604.11632v1 Announce Type: new Abstract: We introduce CARTBENCH, a museum-grounded benchmark for evaluating vision-language models (VLMs) on Chinese artworks beyond short-form recognition and Q
arXiv:2604.11632v1 Announce Type: new Abstract: We introduce CARTBENCH, a museum-grounded benchmark for evaluating vision-language models (VLMs) on Chinese artworks beyond short-form recognition and QA. CARTBENCH comprises four subtasks: CURATORQA for evidence-grounded recognition and reasoning, CATALOGCAPTION for structured four-section expert-style appreciation, REINTERPRET for defensible reinterpretation with expert ratings, and CONNOISSEURPAIRS for diagnostic authenticity discrimination under visually similar confounds. CARTBENCH is built by aligning image-bearing Palace Museum objects from Wikidata with authoritative catalog pages, spanning five art categories across multiple dynasties. Across nine representative VLMs, we find that high overall CURATORQA accuracy can mask sharp drops on hard evidence linking and style-to-period inference; long-form appreciation remains far from expert references; and authenticity-oriented diagnostic discrimination stays near chance, underscoring the difficulty of connoisseur-level reasoning for current models.
Related
- KMMMU: Evaluation of Massive Multi-discipline Multimodal Understanding in Korean Language and Context
- Which Way Does Time Flow? A Psychophysics-Grounded Evaluation for Vision-Language Models
- Disparities In Negation Understanding Across Languages In Vision-Language Models
- Do Vision-Language Models Truly Perform Vision Reasoning? A Rigorous Study of the Modality Gap
- Finding Culture-Sensitive Neurons in Vision-Language Models
Source: arXiv cs.CL | 2026-04-14