Looped Diffusion Transformer

Published
Source
arXiv
Paper number
1141
Field
Computer Vision
arXiv ID
2609.40305

Key points

  • Instead of scaling up model size, the paper gains capacity by looping shared transformer blocks multiple times within a single denoising step.
  • It identifies weak supervision across intermediate loops and unstable attention updates as the reason naive looping degrades quality.
  • Looped-DiT, combining intermediate-loop deep supervision with self-modulating attention, lets a 260M-parameter model surpass a model 6.5x larger.
  • Under a fixed inference budget, increasing loop depth yielded larger gains than adding more denoising steps.
  • Deeper loops progressively corrected mistakes made in earlier loops, showing behavior suggestive of latent reasoning.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)