LongTraceRL: Learning Long-Context Reasoning from Search Agent Trajectories with Rubric Rewards
- Published
- Source
- arXiv
- Paper number
- 278
- Field
- LLMs / NLP
- arXiv ID
- 2605.31584
Key points
- Experiments on five long-context benchmarks with three reasoning LLMs, ranging from 4B to 30B, show that LongTraceRL consistently outperforms strong baselines and encourages comprehensive, evidence-based reasoning.
- Reinforcement learning with verifiable rewards, or RLVR, has shown promise for this task, but prior methods are limited to low-confusion distractors and sparse, outcome-only reward signals, so they fail to supervise intermediate reasoning steps.
- This paper proposes LongTraceRL to address these issues.
Paper links
External research summaries. These are not HDATF publications or measured product results.