Qwen-Image-2.0-RL Technical Report

Published
Source
arXiv
Paper number
514
Field
Computer Vision
arXiv ID
2606.27608

Key points

  • It fine-tunes a VLM with pointwise scoring and chain-of-thought reasoning to build task-specific composite reward models for T2I, including alignment, aesthetics, and portrait quality, and for editing, including instruction following and face ID.
  • It introduces a hybrid CFG strategy in the GRPO-based RL framework to preserve pretrained knowledge and training stability at the same time.
  • On-policy distillation, or OPD, integrates the T2I teacher and the editing teacher into a single student model through trajectory-level velocity matching.
  • OPD achieves better performance than mixed RL, which optimizes multiple rewards at once, because it avoids cross-task interference.
  • It reports consistent improvements on Qwen-Image-Bench with 57.84 points, up 2.61, on the T2I arena with an Elo gain of 78, and on the editing arena with an Elo gain of 93.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)