HarnessOpt-Bench: Evaluating LLMs at Harness Optimization

Published
Source
arXiv
Paper number
840
Field
AI / General
arXiv ID
2608.06301

Key points

  • The authors design a trustworthy execution environment that turns harness optimization into a reproducible evaluation target.
  • They separate the model, harness, and task effects through 111 evaluations across five frontier models and four tasks.
  • Differences between optimization models are larger on average than differences between coding harnesses.
  • Broader modifications are associated with larger improvements, while detailed failure-trace analysis is not correlated with improvement.
  • A native harness is not always better than a shared harness, with 11 wins and 9 losses.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)