KVAE: Family of Tokenizers for Multimodal Generative Models
- Published
- Source
- arXiv
- Paper number
- 843
- Field
- Computer Vision
- arXiv ID
- 2608.05798
Key points
- It designed the KVAE tokenizer family with two compression ratios for video, eightfold compression for images, and 48kHz audio with a 50Hz latent representation.
- In video and image reconstruction and generation evaluations, it reported results matching or exceeding public tokenizers such as Wan-2.2, HunyuanVideo-1.5, and FLUX.2.
- The audio model has 166.9 million parameters, making it smaller than the larger comparison models, while showing preference advantages in multiple domains in human side-by-side listening evaluations.
- It treated diffusability, whether the latent space is easy for a generative model to learn, as a design criterion alongside reconstruction quality, and released the code and models under the MIT License.
- The audio FAD evaluator internally handles only 16–32kHz and therefore cannot measure full-band quality above 16kHz, while model-based aesthetic scores are also subject to evaluator bias.
Paper links
External research summaries. These are not HDATF publications or measured product results.