VCU-Bridge: Hierarchical Visual Connotation Understanding via Semantic Bridging
arXiv:2511.18121v2 Announce Type: replace-cross Abstract: While Multimodal Large Language Models (MLLMs) excel on benchmarks, their processing paradigm differs from the human ability to integrate visu