VespaSeg: A Resource-Aware Ground-then-Segment Pipeline for Referring Expression Segmentation
arXiv:2608.01077v1 Announce Type: new Abstract: Referring expression segmentation requires language conditioned localization and pixel-accurate masks, but monolithic models can be costly to deploy. We