UniJEPA: A Unified Joint-Embedding Predictive Architecture for Task-Agnostic Visual World Modeling

Published
Source
arXiv
Paper number
856
Field
Computer Vision
arXiv ID
2608.07409

Key points

  • It proposes a single objective that jointly learns image-level transformation prediction and video-level temporal prediction in one latent space.
  • End-to-end training is possible without EMA, stop-gradient, or a pretrained encoder, and Gaussian regularization prevents representation collapse in a mathematically grounded way.
  • Photometric prediction learns invariant structure, temporal prediction learns equivariant dynamics, and the balance between them controls the level of abstraction.
  • It achieves competitive or better performance than individual JEPA models on ImageNet, Something-Something-v2, Epic-Kitchens, and control benchmarks.
  • After action-conditioned post-training, it supports zero-shot planning and plans up to tens of times faster than generative world models.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)