UniJEPA: A Unified Joint-Embedding Predictive Architecture for Task-Agnostic Visual World Modeling
- Published
- Source
- arXiv
- Paper number
- 856
- Field
- Computer Vision
- arXiv ID
- 2608.07409
Key points
- It proposes a single objective that jointly learns image-level transformation prediction and video-level temporal prediction in one latent space.
- End-to-end training is possible without EMA, stop-gradient, or a pretrained encoder, and Gaussian regularization prevents representation collapse in a mathematically grounded way.
- Photometric prediction learns invariant structure, temporal prediction learns equivariant dynamics, and the balance between them controls the level of abstraction.
- It achieves competitive or better performance than individual JEPA models on ImageNet, Something-Something-v2, Epic-Kitchens, and control benchmarks.
- After action-conditioned post-training, it supports zero-shot planning and plans up to tens of times faster than generative world models.
Paper links
External research summaries. These are not HDATF publications or measured product results.