Agents

VOS-Agent: The 1st Place Solution for the 8th LSVOS Challenge (MOSEv2 Track)

arXiv:2608.12721v1 Announce Type: new Abstract: Complex video object segmentation requires robust target propagation under severe occlusion, disappearance and reappearance. Although SAM3 provides stro

DGX agentpaper
agentsarxiv-cs-cv

arXiv:2608.12721v1 Announce Type: new Abstract: Complex video object segmentation requires robust target propagation under severe occlusion, disappearance and reappearance. Although SAM3 provides strong promptable mask propagation, a uniform inference path remains unreliable for tiny targets with insufficient visual evidence and semantic-dominated targets whose identities depend on explicit attributes. To this end, we present VOS-Agent, a collaborative framework that retains SAM3 as the shared dense segmentation module and conditionally activates specialized agents according to target characteristics. A Target Perception and Routing Agent assigns each sequence to a regular, tiny, or semantic-dominated route. Tiny targets are supported by a Visual Tracking Agent through confidence-aware box prompts, while semantic-dominated targets are handled by an MLLM-based Semantic Agent through description-guided localization and candidate verification. On the MOSEv2 test set, VOS-Agent achieves 69.82% on the official J&ot{F} metric and ranks first in the MOSEv2 Track of the 8th LSVOS Challenge at ECCV 2026.

Related

Source: arXiv cs.CV | 2026-08-14

Loading related sources…