RREDCoT: Segment-Level Reward Redistribution for Reasoning Models
- Published
- Source
- arXiv
- Paper number
- 333
- Field
- Machine Learning
- arXiv ID
- 2606.06475
Key points
- The method addresses the delayed-reward problem in chain-of-thought reasoning through segment-level reward redistribution.
- We use the model itself to approximate optimal reward redistribution without additional generation.
- This greatly reduces computational overhead compared with Monte Carlo sampling.
- We perform credit assignment by identifying CoT segments that contribute to reaching the correct answer.
- We systematically analyze design choices for CoT segmentation and state-value estimation.
Paper links
External research summaries. These are not HDATF publications or measured product results.