TRACE: Turn-level Reward Assignment via Credit Estimation for Long-Horizon Agents
- Published
- Source
- arXiv
- Paper number
- 633
- Field
- Machine Learning
- arXiv ID
- 2607.13988
Key points
- At each state on a tool-call boundary, it computes the log-probability of the correct answer using a frozen reference model, then uses temporal differences in log-ratio state values as per-action rewards.
- Without a separate critic or process labels, it provides finer-grained turn-level credit than outcome rewards, and values cancel across unnecessary tool calls.
- On BrowseComp-Plus, Qwen3-4B improved from 7.2 to 35.6, and Qwen3-30B-A3B from 8.4 to 42.6.
- By teaching useful intermediate actions in long search processes, it accelerated convergence in pure reinforcement learning and transferred behavior to public-web evaluations.
- Current value estimation is tailored to search with short, known correct answers, so whether the same log-probability criterion is suitable for long code patches or open-ended answers remains unverified.
Paper links
External research summaries. These are not HDATF publications or measured product results.