Video Generative Models as Geometry Learner
- Published
- Source
- arXiv
- Paper number
- 1045
- Field
- Computer Vision
- arXiv ID
- 2608.28549
Key points
- By directly inheriting a video-generation model's existing knowledge of temporal consistency, it generated depth and normals together with images in a single pass.
- With approximately one hundredth as much training data, it matched existing state-of-the-art discriminative models.
- Unlike previous methods that predicted depth and normals separately, it unified them in a single model, saving checkpoint storage and computation costs.
- A single prediction surpassed all generative baselines, and a lightweight ensemble (5×5, 10 seconds) further improved performance.
- The extracted 3D geometry enabled applications such as image relighting, controllable image generation, and 3D mesh reconstruction.
Paper links
External research summaries. These are not HDATF publications or measured product results.