RepFusion: Leveraging Multimodal Priors for Denoising in Representation Space
- Published
- Source
- arXiv
- Paper number
- 417
- Field
- Computer Vision
- arXiv ID
- 2606.14700
Key points
- It extends the MLLM's MLP projector in the RAE latent space to noisy inputs and conditions the DiT with a frozen LLM backbone.
- At similar inference FLOPs, RepFusion's 1.3B DiT outperforms TextEmbed at 8B DiT and Transfusion at 8B joint.
- When switching from VAE to RAE, RepFusion improves by 30 absolute points, from 0.54 to 0.70, which is a larger gain than TextEmbed's 21 percent and Transfusion's 11 percent.
- It performs better when the MLLM is frozen than when it is fine-tuned, which shows that preserving the pretraining prior is important.
- Test-time compute can be used efficiently through stepwise MLLM reconditioning.
- RepFusion-SFT reaches state-of-the-art-level performance with GenEval at 0.87, GenEval2 at 34.9, and DPG at 85.11.
Paper links
External research summaries. These are not HDATF publications or measured product results.