WeChat Vision's universal multimodal embedding family: text, images, videos, visual documents and arbitrarily interleaved inputs mapped to a shared space with flexible output dimensions, in 2B, 4B and 9B variants (weights on HuggingFace, August 26, 2026). Trained in two stages — large-scale multimodal alignment, then refinement with curated data, fine-grained relevance supervision and cross-scale knowledge transfer — and reports leading results on public multimodal-embedding benchmarks.

Model Details

Variants

Name Parameters Notes
WeMM-Embedding-2B 2B —
WeMM-Embedding-4B 4B —
WeMM-Embedding-9B 9B —

Paper

embeddingsmultimodalopen-weight

Related