OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation
- Published
- Source
- arXiv
- Paper number
- 734
- Field
- Computer Vision
- arXiv ID
- 2607.23855
Key points
- Training audio VAE and video VAE separately causes their latent spaces to diverge, making synchronization difficult.
- OmniVAE is an audio-video VAE that finely aligns the two spaces through joint training.
- It matches temporal semantics with segment-level contrastive learning.
- It distills semantic features from a pretrained encoder to improve the trainability of each latent space.
- Both objectives are used only during training, so inference cost does not increase.
- On Verse-Bench, both synchronization, DeSync dropping from 0.884 to 0.570, and audio quality improve.
Paper links
External research summaries. These are not HDATF publications or measured product results.