LeVJEPA: Efficient & Scalable Video Pretraining without the Heuristics

Published
Source
arXiv
Paper number
1029
Field
Computer Vision
arXiv ID
2608.27395

Key points

  • It learned video representations using only a single encoder, without an EMA target encoder, stop-gradient, or predictor.
  • Randomly dropping 95% of tokens actually increased ImageNet accuracy from 33.9% to 47.6%.
  • With the same data and number of epochs, it matched or exceeded V-JEPA 2's performance while reducing total pretraining compute by a factor of 5.6–20.8.
  • At the same compute budget, it outperformed the strongest video baseline by 7.6 points on ImageNet-1K.
  • Training with block-causal attention allows representations to be extended at constant cost as new frames arrive, making it well suited to streaming and world models.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)