An Empirical Study of Training Pixel-Space Text-to-Image Diffusion Models
- Published
- Source
- arXiv
- Paper number
- 923
- Field
- Computer Vision
- arXiv ID
- 2608.16887
Key points
- Experiments that changed only the prediction space while keeping the same Z-Image architecture and corpus of more than 20 billion image-text pairs showed the slow convergence of pixel-space pretraining.
- When run at 1,024 resolution on a single H800 and only the decoder was used, the DiP decoder used 10.09 million parameters and 6.02 milliseconds, and recorded GenEval 0.7545 and DPG 87.54.
- In the noise-scale experiment, 2 worked best with GenEval 0.7545 and DPG 87.54. The theoretically derived value of 8 reached only 0.7316 and 85.64, respectively.
- At 1,024 resolution on a single H800 with 100 model evaluations, Z-Image generation time dropped from 20.12 seconds to 4.56 seconds, and GenEval rose from 0.7510 to 0.7644.
- In the same Z-Image comparison, OneIG fell slightly from 0.566 to 0.564, and LongText from 0.9332 to 0.9252. The validated model families are Z-Image and FLUX2-klein.
Paper links
External research summaries. These are not HDATF publications or measured product results.