UEmbed: Unified Sparse and Dense Multimodal Embeddings
- Published
- Source
- arXiv
- Paper number
- 801
- Field
- Computer Vision
- arXiv ID
- 2608.02583
Key points
- It attaches 16 special tokens and partitions the vocabulary into 16 k-means clusters so that each token owns one chunk, and the number of special tokens N was swept across 2, 4, 8, 16, and 32, with stable performance up to 16 and a clear drop at 32.
- It compresses the vocabulary first, merging tokens that become identical after accent removal, lowercasing, and whitespace normalization, which reduces the vocabulary size from 248,320 to 184,016, and scoring uses the maximum weight among merged tokens.
- UEmbed-9B scores 71.8 dense and 71.0 sparse on MMEB-v2, making it the best dense retrieval model among those trained on public data and a new state-of-the-art for sparse retrieval.
- For UEmbed-2B, using a combined score that sums dense and sparse results raises text from 53.6 to 53.9 and visual documents from 77.0 to 77.5, while natural images and videos see little benefit because they contain less surface vocabulary information.
- On the agentic retrieval benchmark BrowseComp-Plus, sparse mode reduces the number of search rounds. UEmbed-2B drops from 39.19 dense rounds to 32.67 sparse rounds while recall rises from 53.47 to 57.04, and the 9B model keeps the same 49.76 accuracy while reducing rounds from 33.68 to 31.05.
- The authors also state the limits honestly: modern LLMs with large vocabularies can activate odd tokens such as "_alt", and because the training corpus is biased toward English and Chinese, tokens from other languages are much less likely to appear.
Paper links
External research summaries. These are not HDATF publications or measured product results.