Harness-R1: Learning to Edit Executable Runtime Harnesses from Agent Failure Trajectories
- Published
- Source
- arXiv
- Paper number
- 815
- Field
- AI / General
- arXiv ID
- 2608.02276
Key points
- It proposes the first learning-based method that automatically generates executable patches for runtime harnesses from agent failure trajectories.
- A 9B harness engineer is trained only with GRPO while the target agent weights remain frozen, which creates a stable pipeline.
- It improves success rate by 9.3 percentage points on average across WebShop, ALFWorld, and DBBench.
- Even after target-agent fine-tuning, it adds another 5.0 percentage points, which shows that the harness engineer and the agent can co-evolve.
- Transfer experiments on 20 different target settings show an average improvement of 7.06 percentage points, which confirms generalization.
Paper links
External research summaries. These are not HDATF publications or measured product results.