Beyond Pixels: From Video Priors to 4D Worlds
- Published
- Source
- arXiv
- Paper number
- 877
- Field
- Computer Vision
- arXiv ID
- 2608.10744
Key points
- It proposes Latent-to-4D, a framework that connects the VAE latent space of video generation models directly to a 4D decoder.
- It skips the step of generating RGB images, which avoids information loss and error accumulation.
- It is trained on about 1,000 existing reconstruction clips and can be reused across multiple video generators with a single checkpoint.
- On the Text4D-200 benchmark, it improves DINO-F1 by 2.88 to 3.45 points, and on I4D-200 it improves by 5.81 points.
- Human evaluation shows that it outperforms prior methods in geometric validity, completeness, and temporal stability.
Paper links
External research summaries. These are not HDATF publications or measured product results.