Video Generation Models are General-Purpose Vision Learners

Published
Source
arXiv
Paper number
593
Field
Computer Vision
arXiv ID
2607.09024

Key points

  • It proposes video diffusion models as a pretraining paradigm for general-purpose vision recognition, serving a catalytic role for vision analogous to next-token prediction in NLP.
  • It converts iterative denoising into a single forward pass that can control depth, normal, segmentation, pose, keypoints, and more through text instructions.
  • It matches or exceeds specialized models such as D4RT and VGGT-Ω while requiring 7x to 500x less training data.
  • It experimentally shows that diffusion-based pretraining outperforms prior pretraining methods such as V-JEPA and VideoMAE V2.
  • Training on synthetic human videos alone still yields zero-shot generalization to real videos and OOD objects, evidence that world models are internalized.
  • Without dedicated heads or losses, the unified backbone, head, and loss structure lets tasks expand simply by changing the data format.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)