ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning

Published
Source
arXiv
Paper number
829
Field
AI / General
arXiv ID
2608.03972

Key points

  • We prove that expert-model failure trajectories, GNT, are far more effective for iterative learning than failures from itself or from weaker models.
  • GNT includes both a valid prefix close to the correct answer and local error points, making it high-quality material for reflective learning.
  • Reflective-to-Direct Policy Transition gradually shifts capability from hint dependence to solving on its own.
  • It achieves consistent accuracy gains across nine benchmarks, four models from 1.5B to 8B, and four training methods.
  • It cuts response length by more than half while improving accuracy and preventing entropy collapse.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)