SRPO: Self-Reflective Policy Optimization for Long-Horizon Reasoning

Published
Source
arXiv
Paper number
996
Field
AI / General
arXiv ID
2608.23493

Key points

  • It derives self-reflection summaries from completed trajectories and uses them as conditions to generate dense token-level training signals.
  • It runs using only the model itself, without a larger teacher, external critic, or separate reward model, allowing post-training even where access to more capable models is difficult.
  • With Qwen3-8B, it achieved 73.3% on AIME'24 while using only 8% (0.08 times) the computation of scaled-up SFT.
  • It also substantially improved success on long-horizon tasks, reaching 64.7% on WebShop, 76.8% on ALFWorld, and 31.2% on SWE-Bench-Lite.
  • However, reflection does not always work: 42% of failed reflections were vague advice, 35% misdiagnosed the cause, and 23% concerned problems requiring knowledge the model did not possess in the first place.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)