ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning
- Published
- Source
- arXiv
- Paper number
- 829
- Field
- AI / General
- arXiv ID
- 2608.03972
Key points
- We prove that expert-model failure trajectories, GNT, are far more effective for iterative learning than failures from itself or from weaker models.
- GNT includes both a valid prefix close to the correct answer and local error points, making it high-quality material for reflective learning.
- Reflective-to-Direct Policy Transition gradually shifts capability from hint dependence to solving on its own.
- It achieves consistent accuracy gains across nine benchmarks, four models from 1.5B to 8B, and four training methods.
- It cuts response length by more than half while improving accuracy and preventing entropy collapse.
Paper links
External research summaries. These are not HDATF publications or measured product results.