LeVLJEPA: End-to-End Vision-Language Pretraining Without Negatives
- Published
- Source
- arXiv
- Paper number
- 543
- Field
- Computer Vision
- arXiv ID
- 2607.00784
Key points
- This is the first proof that stable large-scale vision-language pretraining is possible with only cross-modal prediction and SIGReg regularization, without negative pairs.
- When used as a VLM backbone, it outperforms both CLIP and SigLIP, with GQA at 44.6%, VQAv2 at 54.2%, and POPE at 66.9% on the Llama-1N setup.
- For semantic segmentation, it provides better dense semantic features than contrastive learning, preserving structured information at the patch-token level.
- In zero-shot classification, contrastive learning still leads, but as a backbone the zero-shot score is actually inversely correlated with performance.
- The entire training recipe removes the need for negatives, temperature, momentum, and teacher-student mechanisms, simplifying training and reducing cost.
- It is well suited to object-centric representation learning and is more robust to backgrounds than contrastive methods.
Paper links
External research summaries. These are not HDATF publications or measured product results.