Cognitive Mismatch in Multimodal Large Language Models for Discrete Symbol Understanding
arXiv:2603.18472v2 Announce Type: replace-cross Abstract: Multimodal large language models (MLLMs) perform strongly on natural images, yet their ability to understand discrete visual symbols remains u