Looped Diffusion Transformer
- Published
- Source
- arXiv
- Paper number
- 1141
- Field
- Computer Vision
- arXiv ID
- 2609.40305
Key points
- Instead of scaling up model size, the paper gains capacity by looping shared transformer blocks multiple times within a single denoising step.
- It identifies weak supervision across intermediate loops and unstable attention updates as the reason naive looping degrades quality.
- Looped-DiT, combining intermediate-loop deep supervision with self-modulating attention, lets a 260M-parameter model surpass a model 6.5x larger.
- Under a fixed inference budget, increasing loop depth yielded larger gains than adding more denoising steps.
- Deeper loops progressively corrected mistakes made in earlier loops, showing behavior suggestive of latent reasoning.
Paper links
External research summaries. These are not HDATF publications or measured product results.