FlowMimic: Mask-free Visual Editing and Generation with Pixel-pair Warped Flow Field for Online Video Editing Data Generation and Modality Mimicry
- Published
- Source
- arXiv
- Paper number
- 675
- Field
- Computer Vision
- arXiv ID
- 2607.18227
Key points
- The core claim is that video editing training does not need truly natural video pairs, only pixel correspondences from the first frame that remain temporally consistent.
- It converts image-editing samples into video-editing samples in real time by turning both the input image and the target image into 3D grids and building a 4D temporally warped flow field over pixel pairs through grid inverse sampling.
- It treats an image as a one-frame video and uses two modality-mimic losses so that T2I mimics the realism of T2V and V2V editing mimics I2I editing, which converges faster.
- It introduces sense-related tasks such as referring expression segmentation, together with edit-region-aware latent-level and attention-level losses, so the model can locate the edit region without a separate MLLM or mask input at inference time.
- It continues training from Wan2.1-T2V-1.3B, and when the learning rate is warmed up from 5e-6 to 1e-5, video-editing behavior appears within a few thousand steps.
- Without using any specially made video-editing data, it still handles a wide range of editing tasks and even shows SAM3-like ability to segment and track objects specified by language in video.
Paper links
External research summaries. These are not HDATF publications or measured product results.