KVAE: Family of Tokenizers for Multimodal Generative Models

Published
Source
arXiv
Paper number
843
Field
Computer Vision
arXiv ID
2608.05798

Key points

  • It designed the KVAE tokenizer family with two compression ratios for video, eightfold compression for images, and 48kHz audio with a 50Hz latent representation.
  • In video and image reconstruction and generation evaluations, it reported results matching or exceeding public tokenizers such as Wan-2.2, HunyuanVideo-1.5, and FLUX.2.
  • The audio model has 166.9 million parameters, making it smaller than the larger comparison models, while showing preference advantages in multiple domains in human side-by-side listening evaluations.
  • It treated diffusability, whether the latent space is easy for a generative model to learn, as a design criterion alongside reconstruction quality, and released the code and models under the MIT License.
  • The audio FAD evaluator internally handles only 16–32kHz and therefore cannot measure full-band quality above 16kHz, while model-based aesthetic scores are also subject to evaluator bias.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)