Harness-R1: Learning to Edit Executable Runtime Harnesses from Agent Failure Trajectories

Published
Source
arXiv
Paper number
815
Field
AI / General
arXiv ID
2608.02276

Key points

  • It proposes the first learning-based method that automatically generates executable patches for runtime harnesses from agent failure trajectories.
  • A 9B harness engineer is trained only with GRPO while the target agent weights remain frozen, which creates a stable pipeline.
  • It improves success rate by 9.3 percentage points on average across WebShop, ALFWorld, and DBBench.
  • Even after target-agent fine-tuning, it adds another 5.0 percentage points, which shows that the harness engineer and the agent can co-evolve.
  • Transfer experiments on 20 different target settings show an average improvement of 7.06 percentage points, which confirms generalization.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)