UniDDT: Unifying Multimodal Understanding and Generation with Decoupled Diffusion Transformer

Published
Source
arXiv
Paper number
434
Field
Computer Vision
arXiv ID
2606.16255

Key points

  • A noisy ViT encoder serves as a unified semantic encoder for both understanding, with clean inputs, and generation, with noisy inputs.
  • A separate diffusion decoder separates text decoding from visual generation, which minimizes modality interference.
  • The model constructs a dual data structure from the same image-text pair so that it can exploit the duality between understanding and generation.
  • It achieves 0.87 on GenEval, 86.9 on DPG, 1699.5 on MME, and 76.5 on SEEDbench, which shows competitive performance on both understanding and generation.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)