Learning Native Reflection in Unified Models with Interleaved Reinforcement Learning

Published
Source
arXiv
Paper number
1135
Field
Computer Vision
arXiv ID
2609.35767

Key points

  • Trained the entire self-reflection loop (diagnose then revise) inside one unified multimodal model with reinforcement learning.
  • Used group-relative advantages over sibling trajectories sharing one initial image so the model compares reflection strategies rather than lucky first draws.
  • Applied a single trajectory-level advantage to update both reflection tokens and image revisions, avoiding the combinatorial blow-up of per-round credit assignment.
  • Gains transferred to benchmarks never used in training: +12.05 on GenEval, +10.97 on WISE, and +3.48 on OneIG-Bench.
  • Found that SFT rollouts already contained correct repairs for 78% of failing images, yet a single trajectory fixed only 20.59%, rising to 64.94% after RL.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)