Local Ai
tencent/WeMM-Embedding 9B/4B/2B
WeMM-Embedding-9B is a universal multimodal embedding model built on Qwen3.5. It accepts text, images, videos, visual documents, and interleaved multimodal inputs, and returns a 4,096-dimensional L2-n
WeMM-Embedding-9B is a universal multimodal embedding model built on Qwen3.5. It accepts text, images, videos, visual documents, and interleaved multimodal inputs, and returns a 4,096-dimensional L2-normalized embedding. Audio input is not supported. https://huggingface.co/tencent/WeMM-Embedding-9B https://huggingface.co/tencent/WeMM-Embedding-4B https://huggingface.co/tencent/WeMM-Embedding-2B https://github.com/Tencent/WeMM-Embedding/blob/main/assets/WeMM_Embedding_tech_report.pdf submitted by /u/jacek2023 [link] [comments]
Related
- tencent/EVIE-Preview-4.5B · Hugging Face
- NEX-N2-mini: 'There is no Pareto frontier. I am Pareto'. This Qwen3.5-MoE fine tune fixed 3.5 and 3.6 overthinking apparently on my tests.
- Meta is about to release a pixel space model (Tuna-2)
Source: r/LocalLLaMA | 2026-08-25