LeVJEPA: Efficient & Scalable Video Pretraining without the Heuristics
- Published
- Source
- arXiv
- Paper number
- 1029
- Field
- Computer Vision
- arXiv ID
- 2608.27395
Key points
- It learned video representations using only a single encoder, without an EMA target encoder, stop-gradient, or predictor.
- Randomly dropping 95% of tokens actually increased ImageNet accuracy from 33.9% to 47.6%.
- With the same data and number of epochs, it matched or exceeded V-JEPA 2's performance while reducing total pretraining compute by a factor of 5.6–20.8.
- At the same compute budget, it outperformed the strongest video baseline by 7.6 points on ImageNet-1K.
- Training with block-causal attention allows representations to be extended at constant cost as new frames arrive, making it well suited to streaming and world models.
Paper links
External research summaries. These are not HDATF publications or measured product results.