RepFusion: Leveraging Multimodal Priors for Denoising in Representation Space

Published
Source
arXiv
Paper number
417
Field
Computer Vision
arXiv ID
2606.14700

Key points

  • It extends the MLLM's MLP projector in the RAE latent space to noisy inputs and conditions the DiT with a frozen LLM backbone.
  • At similar inference FLOPs, RepFusion's 1.3B DiT outperforms TextEmbed at 8B DiT and Transfusion at 8B joint.
  • When switching from VAE to RAE, RepFusion improves by 30 absolute points, from 0.54 to 0.70, which is a larger gain than TextEmbed's 21 percent and Transfusion's 11 percent.
  • It performs better when the MLLM is frozen than when it is fine-tuned, which shows that preserving the pretraining prior is important.
  • Test-time compute can be used efficiently through stepwise MLLM reconditioning.
  • RepFusion-SFT reaches state-of-the-art-level performance with GenEval at 0.87, GenEval2 at 34.9, and DPG at 85.11.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)