Pixel-Space Diffusion Transformers
- Published
- Source
- arXiv
- Paper number
- 673
- Field
- Computer Vision
- arXiv ID
- 2607.17585
Key points
- It points out that latent diffusion has an information bottleneck because the VAE's lossy compression irreversibly removes high-frequency texture and sharp boundaries.
- It identifies the core problem as a mismatch between the objectives of a compressor trained for reconstruction and a diffusion model trained for generation, which makes end-to-end optimization difficult.
- Moving to pixel space enables single-stage training, but it introduces new challenges in noise schedules, loss weighting, token efficiency, and scalable architecture design.
- Recent methods are introduced as x0 prediction approaches that directly predict the clean original image and speed convergence by keeping the training path on the data surface.
- It summarizes the trend toward splitting global structure and fine texture across different components through multi-branch designs, frequency separation, and finer tokenization.
- It suggests that if images, text, and task conditions are placed in a shared token space and trained with a single transformer, the field can move toward unified vision foundation models that do both understanding and generation.
Paper links
External research summaries. These are not HDATF publications or measured product results.