Agents
VOS-Agent: The 1st Place Solution for the 8th LSVOS Challenge (MOSEv2 Track)
arXiv:2608.12721v1 Announce Type: new Abstract: Complex video object segmentation requires robust target propagation under severe occlusion, disappearance and reappearance. Although SAM3 provides stro
arXiv:2608.12721v1 Announce Type: new Abstract: Complex video object segmentation requires robust target propagation under severe occlusion, disappearance and reappearance. Although SAM3 provides strong promptable mask propagation, a uniform inference path remains unreliable for tiny targets with insufficient visual evidence and semantic-dominated targets whose identities depend on explicit attributes. To this end, we present VOS-Agent, a collaborative framework that retains SAM3 as the shared dense segmentation module and conditionally activates specialized agents according to target characteristics. A Target Perception and Routing Agent assigns each sequence to a regular, tiny, or semantic-dominated route. Tiny targets are supported by a Visual Tracking Agent through confidence-aware box prompts, while semantic-dominated targets are handled by an MLLM-based Semantic Agent through description-guided localization and candidate verification. On the MOSEv2 test set, VOS-Agent achieves 69.82% on the official J&ot{F} metric and ranks first in the MOSEv2 Track of the 8th LSVOS Challenge at ECCV 2026.
Related
- Multi-Agent Target-Existence Verification and Learned Mask Geometry Refinement: Winning Report of the MeViS-Text Track at the 8th LSVOS Challenge 2026
- Seg2Track++: Probabilistic Track Validation and Data Association for Multi-Object Tracking and Segmentation
- AgentRVOS for MeViS-Text Track of 5th PVUW Challenge: 3rd Method
Source: arXiv cs.CV | 2026-08-14