WeMM-Embedding
model Your tags
Your notes
WeChat Vision's universal multimodal embedding family: text, images, videos, visual documents and arbitrarily interleaved inputs mapped to a shared space with flexible output dimensions, in 2B, 4B and 9B variants (weights on HuggingFace, August 26, 2026). Trained in two stages — large-scale multimodal alignment, then refinement with curated data, fine-grained relevance supervision and cross-scale knowledge transfer — and reports leading results on public multimodal-embedding benchmarks.
Model Details
Variants
| Name | Parameters | Notes |
|---|---|---|
| WeMM-Embedding-2B | 2B | — |
| WeMM-Embedding-4B | 4B | — |
| WeMM-Embedding-9B | 9B | — |