FiRE: Enhancing MLLMs with Fine-Grained Context Learning for Complex Image Retrieval
arXiv:2607.27959v1 Announce Type: new Abstract: Due to their strong generalizable multimodal processing and reasoning capabilities, Multimodal Large Language Models (MLLMs) have demonstrated significa