Read It Back: Pretrained MLLMs Are Zero-Shot Reward Models for Text-to-Image Generation

Published
Source
arXiv
Paper number
608
Field
Computer Vision
arXiv ID
2607.11886

Key points

  • We use image-conditioned prompt log-likelihood from an MLLM as the reward, which removes the need for preference labels and fine-tuning.
  • Self-SpectraReward is a closed-loop self-improvement scheme in which the understanding branch of a unified model (UMM) evaluates its own generation branch.
  • We validate it at scale with 9 MLLM backbones from 4B to 235B, 4 model families, 3 RL algorithms, and 2 diffusion models.
  • A larger reward MLLM is not always better; reward-policy alignment matters more than reward-model size.
  • Self-SpectraReward outperforms a 235B external model, demonstrating the importance of policy-reward alignment.
  • It consistently outperforms prior MLLM-derived reward methods on five OOD text-to-image benchmarks.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)