Tencent has released WeMM-Embedding-9B, a universal multimodal embedding model built on Qwen3.5, under the Apache License 2.0. The model accepts text, images, videos, visual documents, and interleaved multimodal inputs, returning 4,096-dimensional L2-normalized embeddings; audio input is not supported. It supports Matryoshka embeddings with reduced dimensions and can be served via vLLM 0.27.0 and SGLang 0.5.9. Benchmark results are reported on MMEB-v2 (78 datasets) and MMEB-v3 (190 tasks).
No score is assigned. Sources and their independence are shown in the citation chain below.