Meta-Harness: End-to-End Optimization of Model Harnesses
- Published
- Source
- arXiv
- Paper number
- 131
- Field
- Agents / Benchmarks
- arXiv ID
- 2603.28052
Key points
- The core idea is to treat the harness as executable stateful code rather than a prompt template, and to search the end-to-end logic of retrieval, memory, prompt assembly, and environment handling.
- In the optimization loop, a coding-agent proposer explores an uncompressed filesystem of prior code, scores, prompts, model outputs, tool calls, and traces, then submits a new harness candidate for evaluation.
- In the empirical results, online classification accuracy was 48.6 percent versus 40.9 percent for ACE while using four times fewer context tokens, average accuracy improved by 4.7 points on 200 IMO-level math problems across five held-out models, and TerminalBench-2 results reached 37.6 percent for Claude Haiku 4.5 and 76.4 percent for Claude Opus 4.6 in the alphaXiv report.
- The significance is that raw traces let the proposer debug long-horizon credit-assignment failures, such as when an early memory miss causes a later answer failure that a score-only or summary-only optimizer would miss.
Paper links
External research summaries. These are not HDATF publications or measured product results.