Video Generative Models as Geometry Learner

Published
Source
arXiv
Paper number
1045
Field
Computer Vision
arXiv ID
2608.28549

Key points

  • By directly inheriting a video-generation model's existing knowledge of temporal consistency, it generated depth and normals together with images in a single pass.
  • With approximately one hundredth as much training data, it matched existing state-of-the-art discriminative models.
  • Unlike previous methods that predicted depth and normals separately, it unified them in a single model, saving checkpoint storage and computation costs.
  • A single prediction surpassed all generative baselines, and a lightweight ensemble (5×5, 10 seconds) further improved performance.
  • The extracted 3D geometry enabled applications such as image relighting, controllable image generation, and 3D mesh reconstruction.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)