SRPO: Self-Reflective Policy Optimization for Long-Horizon Reasoning
- Published
- Source
- arXiv
- Paper number
- 996
- Field
- AI / General
- arXiv ID
- 2608.23493
Key points
- It derives self-reflection summaries from completed trajectories and uses them as conditions to generate dense token-level training signals.
- It runs using only the model itself, without a larger teacher, external critic, or separate reward model, allowing post-training even where access to more capable models is difficult.
- With Qwen3-8B, it achieved 73.3% on AIME'24 while using only 8% (0.08 times) the computation of scaled-up SFT.
- It also substantially improved success on long-horizon tasks, reaching 64.7% on WebShop, 76.8% on ALFWorld, and 31.2% on SWE-Bench-Lite.
- However, reflection does not always work: 42% of failed reflections were vague advice, 35% misdiagnosed the cause, and 23% concerned problems requiring knowledge the model did not possess in the first place.
Paper links
External research summaries. These are not HDATF publications or measured product results.