VideoRAE: Taming Video Foundation Models for Generative Modeling via Representation Autoencoders
- Published
- Source
- arXiv
- Paper number
- 634
- Field
- Computer Vision
- arXiv ID
- 2607.14088
Key points
- It compresses features from multiple layers of a frozen video foundation encoder through a lightweight one-dimensional self-attention projector.
- The continuous latent representations can be used in a diffusion Transformer, while discrete tokens produced through multi-codebook quantization can be used in an autoregressive model.
- On UCF-101, the autoregressive and diffusion generators achieved gFVD scores of 40 and 93, respectively, and converged approximately 5× faster than competing autoencoders.
- It can accelerate video-generation training by converting semantically rich frozen video representations into a reconstructable generative latent space.
- Generation results were verified primarily on UCF-101 and in controlled 2B-scale text-to-video experiments, leaving validation on larger models and more varied video domains outstanding.
Paper links
External research summaries. These are not HDATF publications or measured product results.