AutoSaddler: Automatic Harness Optimization with Durable Updates from Agent Execution Traces
- Published
- Source
- arXiv
- Paper number
- 1005
- Field
- AI / General
- arXiv ID
- 2608.23041
Key points
- It formalizes harness optimization as an offline learning problem and iteratively patches the harness in response to failure signals.
- Deep debugging rather than shallow reflection, structured patches rather than arbitrary changes, and generalization checks were central to its effectiveness.
- It achieved 72.3% on GAIA2 with approximately 1,000 task executions, outperforming a competing method that scored 64.6% using approximately 2,800 executions.
- Across all three benchmarks, it improved by +9.0–10.0 percentage points over the base harness and +4.4–7.4 percentage points over the strongest automatic baseline.
- The authors noted that it requires training tasks with ground-truth answers and success/failure signals, making direct use in real-world settings difficult, and that its applicability is limited to independent tasks without persistent state.
Paper links
External research summaries. These are not HDATF publications or measured product results.