Model Releases
Diverse-Intent Multi-Turn Fashion Image Retrieval
arXiv:2607.20291v1 Announce Type: new Abstract: Real-world fashion search involves interactive retrieval across multiple turns. However, existing multi-turn retrieval methods are built on a restrictiv
arXiv:2607.20291v1 Announce Type: new Abstract: Real-world fashion search involves interactive retrieval across multiple turns. However, existing multi-turn retrieval methods are built on a restrictive assumption that every interaction follows the same attribute-editing paradigm, leaving heterogeneous intent transitions unexplored. Moreover, existing approaches often rely on textification to bridge multimodal queries and visual retrieval, which may lose fine-grained visual cues. To address these gaps, we introduce DIM-Fashion, a benchmark of 26K multi-turn sessions constructed from 13 fashion retrieval datasets across 7 tasks, featuring diverse intent transitions and rollback behaviors. We further propose FashionAM, an MLLM-VLP framework that directly aligns multimodal conversational queries with a fashion-oriented gallery embedding space, avoiding intermediate textification. Extensive experiments demonstrate the effectiveness of FashionAM over existing approaches. The dataset and code will be made publicly available upon acceptance.
Related
- CIRCLED: A Multi-turn CIR Dataset with Consistent Dialogues across Domains
- FashionLens: Toward Versatile Fashion Image Retrieval via Task-Adaptive Learning
- Tstars-Tryon 1.0: Robust and Realistic Virtual Try-On for Diverse Fashion Items
Source: arXiv cs.CV | 2026-07-23