Mage-Flow: An Efficient Native-Resolution Foundation Model for Image Generation and Editing

Published
Source
arXiv
Paper number
682
Field
Computer Vision
arXiv ID
2607.19064

Key points

  • At 4B parameters, it matches the performance of competitors in the 6B to 80B range while keeping memory usage down to about 18 GB.
  • The Mage-VAE tokenizer reduces compute per pixel by 12 to 22 times compared with conventional VAEs.
  • CUDA kernel fusion improves training throughput by about 2.5 times.
  • The turbo variant generates a 1024 by 1024 image in 0.59 seconds and performs editing in 1.02 seconds on a single A100.
  • Its cost efficiency makes it a practical open baseline for researchers who want to experiment and iterate.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)