Gemini Embedding 2: A Native Multimodal Embedding Model from Gemini
- Published
- Source
- arXiv
- Paper number
- 243
- Field
- Computer Vision
- arXiv ID
- 2605.27295
Key points
- Gemini-based native multimodal processing maps text, images, video, audio, and interleaved combinations into a single embedding space.
- It is trained with vectors of up to 3,072 dimensions and large-scale multi-task, multi-stage contrastive learning.
- It achieves SOTA with 62.9 R@1 on MSCOCO, 68.8 NDCG@10 on Vatex, 69.9 on MTEB Multilingual, and 84.0 on MTEB Code.
- Compared with traditional late-fusion models such as CLIP and SigLIP, deep modal fusion provides cross-modal synergy.
- It shows strong zero-shot generalization in specialized domains such as astronomy, biology, art, and cooking.
- It delivers out-of-the-box performance that can be used immediately for practical downstream tasks such as RAG, recommendation, and search.
Paper links
External research summaries. These are not HDATF publications or measured product results.