FlexComposer: Unified Video Compositing from Images to Dynamic Footage with Flexible Trajectory Control

Published
Source
arXiv
Paper number
790
Field
Computer Vision
arXiv ID
2607.29627

Key points

  • It proposes a unified video synthesis framework that handles still images and dynamic videos within the same framework.
  • It standardizes heterogeneous inputs with a unified normalized foreground representation that separates the foreground's intrinsic motion from global movement.
  • It uses spatially aware latent injection based on the translation equivariance of the VAE latent space to achieve precise trajectory control without extra trainable modules.
  • It learns lighting and shadows naturally from a hybrid dataset that combines simulation, real movie footage, and generated data, together with a synthetic-to-real curriculum.
  • It outperforms prior state-of-the-art methods in visual quality, temporal consistency, and trajectory adherence.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)