LeVLJEPA: End-to-End Vision-Language Pretraining Without Negatives

Published
Source
arXiv
Paper number
543
Field
Computer Vision
arXiv ID
2607.00784

Key points

  • This is the first proof that stable large-scale vision-language pretraining is possible with only cross-modal prediction and SIGReg regularization, without negative pairs.
  • When used as a VLM backbone, it outperforms both CLIP and SigLIP, with GQA at 44.6%, VQAv2 at 54.2%, and POPE at 66.9% on the Llama-1N setup.
  • For semantic segmentation, it provides better dense semantic features than contrastive learning, preserving structured information at the patch-token level.
  • In zero-shot classification, contrastive learning still leads, but as a backbone the zero-shot score is actually inversely correlated with performance.
  • The entire training recipe removes the need for negatives, temperature, momentum, and teacher-student mechanisms, simplifying training and reducing cost.
  • It is well suited to object-centric representation learning and is more robust to backgrounds than contrastive methods.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)