V-RAE: Rethinking Video Latent Spaces for Generation
- Published
- Source
- arXiv
- Paper number
- 895
- Field
- Computer Vision
- arXiv ID
- 2608.13556
Key points
- Instead of learning the video latent space for pixel reconstruction, it builds it on top of the semantic feature space of a frozen visual backbone.
- On K600 reconstruction, it achieves an rFVD of 2.13, which beats every large video VAE in the evaluation, and it preserves semantic information in the latent space much better, with 89.13 percent classification on UCF101 versus a VAE best of 30.83 percent.
- Under the same training conditions, it also leads in generation quality, measured by gFVD, and converges up to 6 times faster.
- It shows that reconstruction ranking, measured by rFVD, can differ greatly from generation quality ranking, and it proposes tFVD as a metric that better reflects generation suitability.
- On Cityscapes future video prediction, it also outperforms the Wan 2.2 VAE latent space, with gFVD 111.36 versus 144.47.
Paper links
External research summaries. These are not HDATF publications or measured product results.