Beyond Pixels: From Video Priors to 4D Worlds

Published
Source
arXiv
Paper number
877
Field
Computer Vision
arXiv ID
2608.10744

Key points

  • It proposes Latent-to-4D, a framework that connects the VAE latent space of video generation models directly to a 4D decoder.
  • It skips the step of generating RGB images, which avoids information loss and error accumulation.
  • It is trained on about 1,000 existing reconstruction clips and can be reused across multiple video generators with a single checkpoint.
  • On the Text4D-200 benchmark, it improves DINO-F1 by 2.88 to 3.45 points, and on I4D-200 it improves by 5.81 points.
  • Human evaluation shows that it outperforms prior methods in geometric validity, completeness, and temporal stability.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)