Decomposed Vision-Language Alignment for Fine-Grained Open-Vocabulary Segmentation
arXiv:2605.15942v1 Announce Type: cross Abstract: Open-vocabulary segmentation models often struggle to generalize to unseen combinations of object categories and attributes, because fine-grained desc