UniFine: A Unified and Fine-grained Approach for Zero-shot Vision-Language Understanding
arXiv:2307.00862v3 Announce Type: replace-cross Abstract: Vision-language tasks, such as VQA, SNLI-VE, and VCR are challenging because they require the model's reasoning ability to understand the sema