PiD: Fast and High-Resolution Latent Decoding with Pixel Diffusion
- Published
- Source
- arXiv
- Paper number
- 223
- Field
- Computer Vision
- arXiv ID
- 2605.23902
Key points
- The pixel diffusion prior is a 1.3B-parameter Transformer optimized for high-resolution synthesis with a rectified-flow objective.
- The latent adapter is a convolutional network that processes the input latent z and aligns it to the spatial grid of the pixel Transformer, ensuring that the latent preserves global structure and semantic information.
- Sigma-aware gating is the key innovation that controls how much the model relies on the latent condition versus its own learned prior, with the latent influence adjusted by the noise level sigma through a gating function.
Paper links
External research summaries. These are not HDATF publications or measured product results.