Energy-Guided Flow Matching

Published
Source
arXiv
Paper number
844
Field
Computer Vision
arXiv ID
2608.05811

Key points

  • It replaced flow matching's fixed clean-image endpoint with an endpoint that moves from a low-frequency image to the original, making generation proceed from coarse structure to detail.
  • It calculated frequency energy for each image and automatically adjusted heat-diffusion time so that the missing energy is released at the same rate.
  • On ImageNet 256×256, it achieved FID 1.55 after 200 training rounds and 1.45 after 600, improving both quality and training efficiency.
  • It achieved FID 1.58 at 512×512, and GenEval 0.85 and DPG-Bench 83.9 for text-to-image generation, without architectural changes or additional losses.
  • It has not yet been evaluated on video and time-varying signals, joint text-image models, embodied decision making, or the largest recent foundation models.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)