TRACE: Turn-level Reward Assignment via Credit Estimation for Long-Horizon Agents

Published
Source
arXiv
Paper number
633
Field
Machine Learning
arXiv ID
2607.13988

Key points

  • At each state on a tool-call boundary, it computes the log-probability of the correct answer using a frozen reference model, then uses temporal differences in log-ratio state values as per-action rewards.
  • Without a separate critic or process labels, it provides finer-grained turn-level credit than outcome rewards, and values cancel across unnecessary tool calls.
  • On BrowseComp-Plus, Qwen3-4B improved from 7.2 to 35.6, and Qwen3-30B-A3B from 8.4 to 42.6.
  • By teaching useful intermediate actions in long search processes, it accelerated convergence in pure reinforcement learning and transferred behavior to public-web evaluations.
  • Current value estimation is tailored to search with short, known correct answers, so whether the same log-probability criterion is suitable for long code patches or open-ended answers remains unverified.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)