WeMM-Embedding: WeChat Multi-Modal Embedding Technical Report

Published
Source
arXiv
Paper number
1015
Field
Computer Vision
arXiv ID
2608.24053

Key points

  • A single model embeds text, images, video, documents, and mixed inputs, with freely selectable output dimensions from 64 to 2048.
  • The 9B model ranked first among all open-source and commercial models on MMEB-v2 with an overall score of 80.6 as of 2026-08-24.
  • Even the small 2B model surpassed the previously best-performing 8B open-source model, demonstrating its efficiency.
  • It has been deployed in search and recommendations for WeChat Channels, Official Accounts, Moments, and commerce, delivering consistent improvements in 14 A/B tests.
  • Even when reduced to 256 dimensions, it retained 98.7% of its performance at 2,048 dimensions, substantially reducing storage and speed burdens.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)