Read It Back: Pretrained MLLMs Are Zero-Shot Reward Models for Text-to-Image Generation
- Published
- Source
- arXiv
- Paper number
- 608
- Field
- Computer Vision
- arXiv ID
- 2607.11886
Key points
- We use image-conditioned prompt log-likelihood from an MLLM as the reward, which removes the need for preference labels and fine-tuning.
- Self-SpectraReward is a closed-loop self-improvement scheme in which the understanding branch of a unified model (UMM) evaluates its own generation branch.
- We validate it at scale with 9 MLLM backbones from 4B to 235B, 4 model families, 3 RL algorithms, and 2 diffusion models.
- A larger reward MLLM is not always better; reward-policy alignment matters more than reward-model size.
- Self-SpectraReward outperforms a 235B external model, demonstrating the importance of policy-reward alignment.
- It consistently outperforms prior MLLM-derived reward methods on five OOD text-to-image benchmarks.
Paper links
External research summaries. These are not HDATF publications or measured product results.