UniDDT: Unifying Multimodal Understanding and Generation with Decoupled Diffusion Transformer
- Published
- Source
- arXiv
- Paper number
- 434
- Field
- Computer Vision
- arXiv ID
- 2606.16255
Key points
- A noisy ViT encoder serves as a unified semantic encoder for both understanding, with clean inputs, and generation, with noisy inputs.
- A separate diffusion decoder separates text decoding from visual generation, which minimizes modality interference.
- The model constructs a dual data structure from the same image-text pair so that it can exploit the duality between understanding and generation.
- It achieves 0.87 on GenEval, 86.9 on DPG, 1699.5 on MME, and 76.5 on SEEDbench, which shows competitive performance on both understanding and generation.
Paper links
External research summaries. These are not HDATF publications or measured product results.