Gemini Embedding 2: A Native Multimodal Embedding Model from Gemini

Published
Source
arXiv
Paper number
243
Field
Computer Vision
arXiv ID
2605.27295

Key points

  • Gemini-based native multimodal processing maps text, images, video, audio, and interleaved combinations into a single embedding space.
  • It is trained with vectors of up to 3,072 dimensions and large-scale multi-task, multi-stage contrastive learning.
  • It achieves SOTA with 62.9 R@1 on MSCOCO, 68.8 NDCG@10 on Vatex, 69.9 on MTEB Multilingual, and 84.0 on MTEB Code.
  • Compared with traditional late-fusion models such as CLIP and SigLIP, deep modal fusion provides cross-modal synergy.
  • It shows strong zero-shot generalization in specialized domains such as astronomy, biology, art, and cooking.
  • It delivers out-of-the-box performance that can be used immediately for practical downstream tasks such as RAG, recommendation, and search.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)