PiD: Fast and High-Resolution Latent Decoding with Pixel Diffusion

Published
Source
arXiv
Paper number
223
Field
Computer Vision
arXiv ID
2605.23902

Key points

  • The pixel diffusion prior is a 1.3B-parameter Transformer optimized for high-resolution synthesis with a rectified-flow objective.
  • The latent adapter is a convolutional network that processes the input latent z and aligns it to the spatial grid of the pixel Transformer, ensuring that the latent preserves global structure and semantic information.
  • Sigma-aware gating is the key innovation that controls how much the model relies on the latent condition versus its own learned prior, with the latent influence adjusted by the noise level sigma through a gating function.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)