FlowMimic: Mask-free Visual Editing and Generation with Pixel-pair Warped Flow Field for Online Video Editing Data Generation and Modality Mimicry

Published
Source
arXiv
Paper number
675
Field
Computer Vision
arXiv ID
2607.18227

Key points

  • The core claim is that video editing training does not need truly natural video pairs, only pixel correspondences from the first frame that remain temporally consistent.
  • It converts image-editing samples into video-editing samples in real time by turning both the input image and the target image into 3D grids and building a 4D temporally warped flow field over pixel pairs through grid inverse sampling.
  • It treats an image as a one-frame video and uses two modality-mimic losses so that T2I mimics the realism of T2V and V2V editing mimics I2I editing, which converges faster.
  • It introduces sense-related tasks such as referring expression segmentation, together with edit-region-aware latent-level and attention-level losses, so the model can locate the edit region without a separate MLLM or mask input at inference time.
  • It continues training from Wan2.1-T2V-1.3B, and when the learning rate is warmed up from 5e-6 to 1e-5, video-editing behavior appears within a few thousand steps.
  • Without using any specially made video-editing data, it still handles a wide range of editing tasks and even shows SAM3-like ability to segment and track objects specified by language in video.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)