LongTraceRL: Learning Long-Context Reasoning from Search Agent Trajectories with Rubric Rewards

Published
Source
arXiv
Paper number
278
Field
LLMs / NLP
arXiv ID
2605.31584

Key points

  • Experiments on five long-context benchmarks with three reasoning LLMs, ranging from 4B to 30B, show that LongTraceRL consistently outperforms strong baselines and encourages comprehensive, evidence-based reasoning.
  • Reinforcement learning with verifiable rewards, or RLVR, has shown promise for this task, but prior methods are limited to low-confusion distractors and sparse, outcome-only reward signals, so they fail to supervise intermediate reasoning steps.
  • This paper proposes LongTraceRL to address these issues.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)