Learning Native Reflection in Unified Models with Interleaved Reinforcement Learning
- Published
- Source
- arXiv
- Paper number
- 1135
- Field
- Computer Vision
- arXiv ID
- 2609.35767
Key points
- Trained the entire self-reflection loop (diagnose then revise) inside one unified multimodal model with reinforcement learning.
- Used group-relative advantages over sibling trajectories sharing one initial image so the model compares reflection strategies rather than lucky first draws.
- Applied a single trajectory-level advantage to update both reflection tokens and image revisions, avoiding the combinatorial blow-up of per-round credit assignment.
- Gains transferred to benchmarks never used in training: +12.05 on GenEval, +10.97 on WISE, and +3.48 on OneIG-Bench.
- Found that SFT rollouts already contained correct repairs for 78% of failing images, yet a single trajectory fixed only 20.59%, rising to 64.94% after RL.
Paper links
External research summaries. These are not HDATF publications or measured product results.