TRIAGE: Role-Typed Credit Assignment for Agentic Reinforcement Learning

Published
Source
arXiv
Paper number
537
Field
Machine Learning
arXiv ID
2606.32017

Key points

  • A core problem in agent RL is that GRPO applies trajectory success or failure to every action equally, which punishes useful exploration in failed rollouts and rewards regression in successful rollouts.
  • It introduces a four-role taxonomy, Decisive, Exploration, No-progress, and Regression, and applies differentiated credit rules by role.
  • It proves that role labels alone can support MSE-optimal segment correction and always reduce advantage estimation error within the judge's trusted range.
  • On Qwen3-1.7B, it improves ALFWorld by 18.4 points and WebShop by 7.9 points, while reducing environment turns per completed rollout by 10.4% to 14.8%.
  • Regression suppression is the main driver of the gains, with exploration bonuses adding only 0.6 to 1.7 points as a stable auxiliary benefit.
  • It outperforms scalar process rewards and outcome-supervised value baselines, so the key benefit comes from role typing itself rather than simply adding dense rewards.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)