VideoRAE: Taming Video Foundation Models for Generative Modeling via Representation Autoencoders

Published
Source
arXiv
Paper number
634
Field
Computer Vision
arXiv ID
2607.14088

Key points

  • It compresses features from multiple layers of a frozen video foundation encoder through a lightweight one-dimensional self-attention projector.
  • The continuous latent representations can be used in a diffusion Transformer, while discrete tokens produced through multi-codebook quantization can be used in an autoregressive model.
  • On UCF-101, the autoregressive and diffusion generators achieved gFVD scores of 40 and 93, respectively, and converged approximately 5× faster than competing autoencoders.
  • It can accelerate video-generation training by converting semantically rich frozen video representations into a reconstructable generative latent space.
  • Generation results were verified primarily on UCF-101 and in controlled 2B-scale text-to-video experiments, leaving validation on larger models and more varied video domains outstanding.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)