OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation

Published
Source
arXiv
Paper number
734
Field
Computer Vision
arXiv ID
2607.23855

Key points

  • Training audio VAE and video VAE separately causes their latent spaces to diverge, making synchronization difficult.
  • OmniVAE is an audio-video VAE that finely aligns the two spaces through joint training.
  • It matches temporal semantics with segment-level contrastive learning.
  • It distills semantic features from a pretrained encoder to improve the trainability of each latent space.
  • Both objectives are used only during training, so inference cost does not increase.
  • On Verse-Bench, both synchronization, DeSync dropping from 0.884 to 0.570, and audio quality improve.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)