Qwen-Image-2.0-RL Technical Report
- Published
- Source
- arXiv
- Paper number
- 514
- Field
- Computer Vision
- arXiv ID
- 2606.27608
Key points
- It fine-tunes a VLM with pointwise scoring and chain-of-thought reasoning to build task-specific composite reward models for T2I, including alignment, aesthetics, and portrait quality, and for editing, including instruction following and face ID.
- It introduces a hybrid CFG strategy in the GRPO-based RL framework to preserve pretrained knowledge and training stability at the same time.
- On-policy distillation, or OPD, integrates the T2I teacher and the editing teacher into a single student model through trajectory-level velocity matching.
- OPD achieves better performance than mixed RL, which optimizes multiple rewards at once, because it avoids cross-task interference.
- It reports consistent improvements on Qwen-Image-Bench with 57.84 points, up 2.61, on the T2I arena with an Elo gain of 78, and on the editing arena with an Elo gain of 93.
Paper links
External research summaries. These are not HDATF publications or measured product results.