HarnessOpt-Bench: Evaluating LLMs at Harness Optimization
- Published
- Source
- arXiv
- Paper number
- 840
- Field
- AI / General
- arXiv ID
- 2608.06301
Key points
- The authors design a trustworthy execution environment that turns harness optimization into a reproducible evaluation target.
- They separate the model, harness, and task effects through 111 evaluations across five frontier models and four tasks.
- Differences between optimization models are larger on average than differences between coding harnesses.
- Broader modifications are associated with larger improvements, while detailed failure-trace analysis is not correlated with improvement.
- A native harness is not always better than a shared harness, with 11 wins and 9 losses.
Paper links
External research summaries. These are not HDATF publications or measured product results.