Model Releases
Illuminating Visual Identity in Universal Multimodal Embeddings
arXiv:2608.01794v1 Announce Type: cross Abstract: Universal Multimodal Embeddings (UMEs) aim to unify various modalities and tasks into a shared representation space. In recent years, this field has w
arXiv:2608.01794v1 Announce Type: cross Abstract: Universal Multimodal Embeddings (UMEs) aim to unify various modalities and tasks into a shared representation space. In recent years, this field has witnessed substantial progress driven by the development of Multimodal Large Language Models (MLLMs). However, a crucial capability, visual identity discrimination, remains underexplored in existing UME methods, despite its critical role in a wide range of tasks, including instance retrieval, re-identification, and identity preservation in AI-generated content. To bridge this gap, we propose a unified formulation for visual identity discrimination~(VisID) and introduce extbf{MVEB} (extbf{M}ultimodal extbf{V}isual Identity extbf{E}mbedding extbf{B}enchmark), a large-scale benchmark curated from both real-world and synthetic datasets to support evaluation and training. Furthermore, we present a simple yet effective learning framework that jointly optimizes general multimodal and visual identity representations through a carefully designed identity-aware sampling mechanism. Extensive experiments demonstrate that our approach successfully endows UMEs with strong identity discrimination capability and maintains competitive general multimodal performance. We believe this work not only illuminates a critical yet neglected capability, but also takes a step toward more holistic universal multimodal embeddings. Code and data are available at href{https://chrisclear3.github.io/MVEB}{MVEB}.
Related
- Think When Needed: Adaptive Reasoning-Driven Multimodal Embeddings with a Dual-LoRA Architecture
- How Far Are We from Generating Missing Modalities with Foundation Models?
- CRAM: Centroid-Routing and Adaptive MoE for Multimodal Continual Instruction Tuning
Source: arXiv cs.CL | 2026-08-04