Meta-Harness: End-to-End Optimization of Model Harnesses

Published
Source
arXiv
Paper number
131
Field
Agents / Benchmarks
arXiv ID
2603.28052

Key points

  • The core idea is to treat the harness as executable stateful code rather than a prompt template, and to search the end-to-end logic of retrieval, memory, prompt assembly, and environment handling.
  • In the optimization loop, a coding-agent proposer explores an uncompressed filesystem of prior code, scores, prompts, model outputs, tool calls, and traces, then submits a new harness candidate for evaluation.
  • In the empirical results, online classification accuracy was 48.6 percent versus 40.9 percent for ACE while using four times fewer context tokens, average accuracy improved by 4.7 points on 200 IMO-level math problems across five held-out models, and TerminalBench-2 results reached 37.6 percent for Claude Haiku 4.5 and 76.4 percent for Claude Opus 4.6 in the alphaXiv report.
  • The significance is that raw traces let the proposer debug long-horizon credit-assignment failures, such as when an early memory miss causes a later answer failure that a score-only or summary-only optimizer would miss.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)