Grid2Matrix: Revealing Digital Agnosia in Vision-Language Models
DGX agentarXiv:2604.09687v1 Announce Type: cross Abstract: Vision-Language Models (VLMs) excel on many multimodal reasoning benchmarks, but these evaluations often do not require an exhaustive readout of the i