Video Generation Models are General-Purpose Vision Learners
- Published
- Source
- arXiv
- Paper number
- 593
- Field
- Computer Vision
- arXiv ID
- 2607.09024
Key points
- It proposes video diffusion models as a pretraining paradigm for general-purpose vision recognition, serving a catalytic role for vision analogous to next-token prediction in NLP.
- It converts iterative denoising into a single forward pass that can control depth, normal, segmentation, pose, keypoints, and more through text instructions.
- It matches or exceeds specialized models such as D4RT and VGGT-Ω while requiring 7x to 500x less training data.
- It experimentally shows that diffusion-based pretraining outperforms prior pretraining methods such as V-JEPA and VideoMAE V2.
- Training on synthetic human videos alone still yields zero-shot generalization to real videos and OOD objects, evidence that world models are internalized.
- Without dedicated heads or losses, the unified backbone, head, and loss structure lets tasks expand simply by changing the data format.
Paper links
External research summaries. These are not HDATF publications or measured product results.