AutoSaddler: Automatic Harness Optimization with Durable Updates from Agent Execution Traces

Published
Source
arXiv
Paper number
1005
Field
AI / General
arXiv ID
2608.23041

Key points

  • It formalizes harness optimization as an offline learning problem and iteratively patches the harness in response to failure signals.
  • Deep debugging rather than shallow reflection, structured patches rather than arbitrary changes, and generalization checks were central to its effectiveness.
  • It achieved 72.3% on GAIA2 with approximately 1,000 task executions, outperforming a competing method that scored 64.6% using approximately 2,800 executions.
  • Across all three benchmarks, it improved by +9.0–10.0 percentage points over the base harness and +4.4–7.4 percentage points over the strongest automatic baseline.
  • The authors noted that it requires training tasks with ground-truth answers and success/failure signals, making direct use in real-world settings difficult, and that its applicability is limited to independent tasks without persistent state.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)