WeMM-Embedding: WeChat Multi-Modal Embedding Technical Report
- Published
- Source
- arXiv
- Paper number
- 1015
- Field
- Computer Vision
- arXiv ID
- 2608.24053
Key points
- A single model embeds text, images, video, documents, and mixed inputs, with freely selectable output dimensions from 64 to 2048.
- The 9B model ranked first among all open-source and commercial models on MMEB-v2 with an overall score of 80.6 as of 2026-08-24.
- Even the small 2B model surpassed the previously best-performing 8B open-source model, demonstrating its efficiency.
- It has been deployed in search and recommendations for WeChat Channels, Official Accounts, Moments, and commerce, delivering consistent improvements in 14 A/B tests.
- Even when reduced to 256 dimensions, it retained 98.7% of its performance at 2,048 dimensions, substantially reducing storage and speed burdens.
Paper links
External research summaries. These are not HDATF publications or measured product results.